Patentable/Patents/US-20260195846-A1
US-20260195846-A1

Cache-Optimized Warp Engines

PublishedJuly 9, 2026
Assigneenot available in USPTO data we have
Technical Abstract

This document describes systems and techniques directed at cache-optimized warp engines. In aspects, a computing device having a first cache memory, a second cache memory, a system memory, and a cache-optimized warp engine is configured to receive a warped input image comprising a plurality of pixel information. Following a coordinate sequence, the cache-optimized warp engine scans the plurality of pixel information and loads a first portion of the pixel information into the first cache memory and a second portion of the pixel information into the second cache memory. Based on the first and second portions of the plurality of pixel information, the cache-optimized warp engine determines first and second portions, respectively, of a plurality of pixel information of a corrected output image. The cache-optimized warp engine stores the first and second portions of the pixel information of the corrected output image as the output image in the system memory.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

receiving, at a computing device having a cache-optimized warp engine, a first cache memory, and a second cache memory, a warped input image comprising a plurality of pixel information; dividing, by the cache-optimized warp engine, the warped input image comprising the plurality of pixel information into N rows; bottom to top and left to right; or bottom to top and right to left; and an intra-row coordinate sequence that is at least one of: an inter-row coordinate sequence that is top to bottom; or top to bottom and left to right; or top to bottom and right to left; and an intra-row coordinate sequence that is at least one of: an inter-row coordinate sequence that is bottom to top; scanning, by the cache-optimized warp engine and following a coordinate sequence including the N rows, the plurality of pixel information of the warped input image, the coordinate sequence comprising at least one of: loading, by the cache-optimized warp engine, a first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device; determining, by the cache-optimized warp engine and based on the first portion of the plurality of pixel information of the warped input image stored in the first cache memory, a first corrected portion of a plurality of pixel information of a corrected output image; loading, by the cache-optimized warp engine, a second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device; determining, by the cache-optimized warp engine and based on the second portion of the plurality of pixel information of the warped input image stored in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image; and storing, by the cache-optimized warp engine, the first portion and the second portion of the plurality of pixel information of the corrected output image as the output image in a system memory of or associated with the computing device. . A method comprising:

2

claim 1 the warped input image is W pixels wide and T pixels tall; W is an integer greater than one; and T is an integer greater than one. . The method of, wherein:

3

claim 2 each row of the N rows is W pixels wide; and each row of the N rows is an integer division of T pixels tall. . The method of, wherein:

4

claim 1 . The method of, wherein the coordinate sequence comprises an intra-row coordinate sequence that is bottom to top and left to right and an inter-row coordinate sequence that is top to bottom.

5

claim 1 the first cache is an L1 cache of the cache-optimized warp engine; and the second cache is an L2 cache of the cache-optimized warp engine; or the second cache is a system-level cache of a system-on-a-chip of the computing device. . The method of, wherein:

6

claim 1 determining the first portion of the plurality of pixel information of the corrected output image uses an arbitrary transform function; or determining the second portion of the plurality of pixel information of the corrected output image uses the arbitrary transform function. . The method of, wherein at least one of:

7

claim 1 piecewise interpolation; linear interpolation; polynomial interpolation; spline interpolation; or mimetic interpolation. . The method of, wherein determining the first or second portion of the plurality of pixel information of the corrected output image uses at least one of:

8

claim 1 lens distortion; geographic distortion; motion compensation; electronic image stabilization; rolling shutter correction; or camera calibration. . The method of, wherein determining the first or second portion of the plurality of pixel information of the corrected output image uses at least one of:

9

claim 1 scanning the plurality of pixel information of the warped input image includes scanning a plurality of memory access units; loading the first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units; or loading the second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units. . The method of, wherein at least one of:

10

claim 9 the at least one memory access unit is L pixels long and S pixels wide; L is an integer greater than one; and S is an integer greater than one. . The method of, wherein:

11

claim 10 . The method of, wherein S is less than L.

12

claim 1 a pixel of the warped input image comprises M bits; and an integer greater than or equal to one; four bits; eight bits; 16 bits; 32 bits; or 64 bits. M is at least one of: . The method of, wherein:

13

claim 1 location information; gamma information; or color information; or the plurality of pixel information of the warped input image comprises: location information; gamma information; or color information. the plurality of pixel information of the corrected output image comprises: . The method of, wherein at least one of:

14

a first cache memory; a second cache memory; system memory; at least one processor; and receive a warped input image comprising a plurality of pixel information; divide the warped input image comprising the plurality of pixel information into N rows; bottom to top and left to right; or bottom to top and right to left; and an inter-row coordinate sequence that is top to bottom; or an intra-row coordinate sequence that is at least one of: top to bottom and left to right; or top to bottom and right to left; and an inter-row coordinate sequence that is bottom to top; an intra-row coordinate sequence that is at least one of: scan, following a coordinate sequence including the N rows, the plurality of pixel information of the warped input image, the coordinate sequence comprising at least one of: load a first portion of the plurality of pixel information of the warped input image into the first cache memory; determine, based on the first portion of the plurality of pixel information of the warped input image stored in the first cache memory, a first corrected portion of a plurality of pixel information of a corrected output image; load a second portion of the plurality of pixel information of the warped input image into the second cache memory; determine, based on the second portion of the plurality of pixel information of the warped input image stored in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image; and store the first portion and the second portion of the plurality of pixel information of the corrected output image as the output image in the system memory. computer-readable media storing instructions that, when executed by the at least one processor, cause the at least one processor to: . A computing device comprising:

15

receive, at a computing device having a first cache memory and a second cache memory, a warped input image comprising a plurality of pixel information; divide the warped input image comprising the plurality of pixel information into N rows; bottom to top and left to right; or bottom to top and right to left; and an inter-row coordinate sequence that is top to bottom; or an intra-row coordinate sequence that is at least one of: top to bottom and left to right; or top to bottom and right to left; and an inter-row coordinate sequence that is bottom to top; an intra-row coordinate sequence that is at least one of: scan, following a coordinate sequence including the N rows, the plurality of pixel information of the warped input image, the coordinate sequence comprising at least one of: load a first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device; determine, based on the first portion of the plurality of pixel information of the warped input image stored in the first cache memory, a first corrected portion of a plurality of pixel information of a corrected output image; load a second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device; determine, based on the second portion of the plurality of pixel information of the warped input image stored in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image; and store the first portion and the second portion of the plurality of pixel information of the corrected output image as the output image in a system memory of or associated with the computing device. . A computer-readable media comprising instructions that, when executed by at least one processor, cause the at least one processor to:

16

claim 14 the warped input image is W pixels wide and T pixels tall; W is an integer greater than one; and T is an integer greater than one. . The computing device of, wherein:

17

claim 16 each row of the N rows is W pixels wide; and each row of the N rows is an integer division of T pixels tall. . The computing device of, wherein:

18

claim 14 . The computing device of, wherein the coordinate sequence comprises an intra-row coordinate sequence that is bottom to top and left to right and an inter-row coordinate sequence that is top to bottom.

19

claim 14 the first cache is an L1 cache; and the second cache is an L2 cache; or the second cache is a system-level cache of a system-on-a-chip of the computing device. . The computing device of, wherein:

20

claim 14 the scan of the plurality of pixel information of the warped input image includes scanning a plurality of memory access units; the load of the first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units; or the load of the second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units. . The computing device of, wherein at least one of:

Detailed Description

Complete technical specification and implementation details from the patent document.

Many computing devices may include a camera for taking photographs and videos. Mobile computing devices (e.g., smartphones, tablets) are often smaller in physical size and may not have space for large lenses or image sensors, which are important for gathering light for better images or videos. Accordingly, many mobile computing devices may include an image-processing unit, or a central-processing unit, configured to provide computational photography features (e.g., light enhancements, blur reduction, distortion correction). Some mobile computing devices may include a warp engine configured to enhance images or videos through lens distortion correction, motion compensation, electronic image stabilization, rolling shutter correction, and so forth. Many warp engines may utilize cache memories to provide these enhancements.

Unfortunately, however, some warp engines may utilize cache memories in an unoptimized fashion through tile-based image processing, which may include dividing a warped input image into tiles (e.g., a grid of portions) that are smaller than the warped input image. The warp engines may scan pixels within a tile by following an unoptimized coordinate sequence. Scanning the pixels may include loading pixel information (e.g., location, color, gamma) associated with the pixels from a system memory (e.g., random-access memory (RAM)) of a computing device. Portions of the warped input image may be loaded into a cache memory for processing subsequent pixels. However, by utilizing the unoptimized coordinate sequence, the cache memory may be larger than necessary, which increases a size and reduces power efficiency of the computing device.

This document describes systems and techniques directed at cache-optimized warp engines (COWEs). In aspects, a computing device having a first cache memory, a second cache memory, a system memory, and a COWE is configured to receive a warped input image comprising a plurality of pixel information. The COWE may receive the warped input image from a camera or the system memory of the computing device. The first cache memory may be a level 1 (L1) cache and include access times that are 100 times faster than that of the system memory. The second cache memory may be a level 2 (L2) cache or a system-level cache (SLC) and include access times that are 10 times faster than that of the system memory.

Following a coordinate sequence, the COWE may scan the plurality of pixel information and load a first portion of the pixel information into the first cache memory and a second portion of the pixel information into the second cache memory. The COWE may scan the plurality of pixel information by loading a memory access unit (MAU) of pixel information from the system memory into the first cache memory or the second cache memory. The first portion of the pixel information may include MAUs having pixel information that may be accessed more frequently by the COWE than the second portion of the pixel information stored in the second cache memory.

Based on the first and second portions of the plurality of pixel information, the COWE may determine first and second portions, respectively, of a plurality of pixel information of a corrected output image. The COWE may utilize a transformation model (e.g., translation, rotation) to first determine locations of the first and second portions of the corrected output image. The COWE may then interpolate pixel information (e.g., color, gamma) of the first and second portions of the corrected output image using any one of a variety of interpolation methods (e.g., bicubic interpolation).

The COWE may store the first and second portions of the pixel information of the corrected output image as an output image in the system memory. The first portion of the pixel information of the corrected output image may further include subsequent first portions of the corrected output image that the COWE determines based on subsequent first portions of the warped input image. Similarly, the second portion of the pixel information of the corrected output image may further include subsequent second portions of the corrected output image that the COWE determines based on subsequent second portions of the warped input image.

The details of one or more implementations are set forth in the accompanying Drawings and the following Detailed Description. Other features and advantages will be apparent from the Detailed Description, the Drawings, and the Claims. This Summary is provided to introduce subject matter that is further described in the Detailed Description. Accordingly, a reader should not consider the Summary to describe essential features or to threshold the scope of the claimed subject matter.

Computing devices (e.g., smartphones, tablets, computers) often include a camera for capturing images and videos. Larger computing devices (e.g., dedicated cameras) may include large lenses and large image sensors capable of capturing more light, which is important for capturing higher-quality images and videos. Smaller computing devices (e.g., smartphones) may not have sufficient physical space to include large lenses or image sensors and, thus, may take advantage of computational photography. In aspects, computational photography may refer to digital image or video capture and processing techniques that use digital computation rather than optical processes. These computing devices may leverage various on-board components (e.g., cameras, processors) and leverage additional sensor data to digitally compute an image. As an example, a computing device may include at least one cache memory, a system memory (e.g., dynamic random-access memory (DRAM)), and a warp engine configured to process a warped (e.g., distorted) input image into a corrected (e.g., undistorted) output image. The warp engine may access the warped input image from the system memory of the computing device, load portions of the warped input image into the cache memory, and process, utilizing the portions of the warped input image in the cache memory and other portions in the system memory, the warped input image into a corrected output image. The portions of the warped input image in the cache memory may be accessed more quickly than the other portions in the system memory, and thus the warped input image may be processed more quickly into the corrected output image.

However, cache memory can be expensive in terms of physical space and power consumption. Accordingly, warp engines that are not optimized to utilize as little cache as possible may result in physically larger and inefficient computing devices. This can be especially problematic in mobile computing devices that may be powered by batteries. For example, some warp engines may process a warped input image by dividing the warped input image into tile lines that include multiple tiles (e.g., a portion of an image). The warp engine may process the warped input image on a tile-by-tile basis by scanning pixels within a tile along a scan line that progresses from left to right, where a tile includes multiple scan lines from top to bottom. The warp engine may scan the pixels by accessing pixel information associated with pixels of a scan line from the system memory of the computing device and loading one or more scan lines of the pixel information into the cache memory. The pixel information may be a number of bits (e.g., four bits, eight bits) and include location information, color information, and/or gamma information for a given pixel. The warp engine may interpolate the pixel information to produce a corrected output pixel.

For inter-tile pixels (e.g., pixels at a border between horizontally adjacent tiles), same pixel information may be used to process a right-most pixel of a current tile and a left-most pixel of a next tile that is on a right of the current tile. For example, the warp engine may load pixel information associated with the right-most pixel of the current tile from system memory of a computing device into a first cache memory of the computing device. The pixel information may include location, color, and gamma information for the right-most pixel and proximate pixels (e.g., two, three, or four pixels in each of four cardinal directions) of the current tile. The warp engine may determine a location of a right-most corrected output pixel of a corrected output tile of a corrected output image using an arbitrary transformation function (e.g., rotation, translation, stretching). After determining the location of the right-most corrected output pixel, the warp engine may interpolate (e.g., using linear interpolation) the color and gamma information to determine color and gamma information associated with the right-most corrected output pixel of the corrected output tile of the corrected output image.

When the warp engine processes the left-most pixel of the next tile, the warp engine may access the pixel information from the first cache memory. Using the pixel information in the first cache memory, the warp engine may determine (e.g., using bicubic interpolation) pixel information (e.g., color, gamma) for a left-most corrected output pixel of a next output tile of the corrected output image. Unfortunately, however, the warp engine may not access location information for the left-most pixel of the next tile prior to loading a scan line of pixel information, so in some implementations the warp engine may store one or more scan lines of pixel information in the first cache memory. Thus, the first cache memory must be at least as large (e.g., in bits, in bytes) as the one or more scan lines of pixel information.

Alternatively, cache-optimized warp engines (COWEs) may process pixel information using rows (e.g., tile lines) rather than tiles. A row may include a portion of a warped input image that is as wide as the warped input image (e.g., in number of pixels) and as tall as an integer division (e.g., a tenth) of the warped input image. Said differently, COWEs may process warped input images on a row-by-row basis by scanning pixels following an intra-row coordinate sequence that is from bottom to top and left to right within a row. Scanning pixels may include fetching pixel information (e.g., color, gamma) stored in a system memory or cache (e.g., if previously fetched) of a computing device. COWEs may then follow the intra-row coordinate sequence for subsequent rows (e.g., rows that are below, rows that are above).

This document describes systems and techniques directed at COWEs. The disclosed systems and techniques may address shortcomings of cache-unoptimized warp engines that may increase sizes and decrease battery lives of mobile computing devices. The conflict between these shortcomings may be addressed by the disclosed systems and techniques, which may provide COWEs, reduce cache sizes, and improve battery life of mobile computing devices.

The following discussion describes operating environments, techniques that may be employed in the operating environments, example devices, and example methods. Although systems and techniques for cache-optimized warp engines are described, it is to be understood that the subject of the appended Claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations, reference to which is made by way of example only.

1 FIG. 100 102 104 106 108 110 112 104 110 106 110 104 106 110 102 110 102 illustrates an example environmentof a computing devicethat includes a first cache memory, a second cache memory, system memory, a processor, and a display. The first cache memorymay be a level 1 (L1) cache associated with the processor. The second cache memorymay be a level 2 (L2) cache associated with the processor. Alternatively, at least one of the first cache memoryor the second cache memorymay be a system-level cache (SLC) associated with the processor, a motherboard (not illustrated), or another appropriate printed circuit board (PCB) with the computing device. The SLC may be a shared cache between the processorand various peripherals or components (e.g., displays, graphics processing units (GPUs), audio codecs) of the computing device.

108 108 108 102 The system memorymay be realized as any one of a variety of volatile memories, including random-access memory (RAM), dynamic random-access memory (DRAM), synchronized dynamic random-access memory (SDRAM), or the like. Alternatively or additionally, the system memorymay be realized as any one of a variety of non-volatile memories, including flash memory (e.g., solid-state drives (SSDs)), read-only memory (ROM), magnetic computer storage devices (e.g., hard disk drives (HDDs), floppy disks, magnetic tape), optical discs (e.g., compact discs (CDs), digital video discs (DVDs)), and so forth. The system memorymay be operably coupled (e.g., electrically, physically, optically) with one or more components of the computing device.

110 110 104 106 108 110 108 The processormay be realized as any one of a variety of single-core or multi-core processors, including central processing units (CPUs), graphics processing units (GPUs), arithmetic logic units (ALUs), reduced instruction set computer (RISC) microprocessors, advanced RISC machine (ARM) microprocessors, and so forth. The processormay utilize the first cache memory, the second cache memory, and the system memoryto provide some or all of the features described herein. The processormay do so by executing computer-readable instructions stored on the system memory.

112 112 112 102 102 112 102 The displaymay be realized as any one of a variety of displays, including a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, an active-matrix organic light-emitting diode (AMOLED) display, a twisted nematic (TN) display, an in-plane switching (IPS) display, and so forth. Further, the displaymay be a display module that includes a stack of layers, including a display, a touchscreen, a digitizer, a display driver integrated circuit (DDIC), a protective cover layer (e.g., glass, plastic), and so forth. The displaymay be included in a housing of the computing deviceor altogether separate from the computing device. The displaymay be used to show content (e.g., images, videos) and/or a graphical user interface (GUI) to users of the computing device.

1 FIG. 102 114 114 114 110 108 114 114 104 106 108 110 further illustrates that the computing deviceincludes a cache-optimized warp engine(COWE). The COWEmay be realized as a standalone peripheral or component (e.g., an image processing unit (IPU)) having dedicated cache memory, or alternatively and as described above, the processormay execute computer-readable instructions stored on the system memoryto provide the COWE(e.g., as a software component). The COWEmay utilize the first cache memory, the second cache memory, the system memory, and the processorto process a warped input image into a corrected output image.

114 114 114 114 An input image may be warped for any one of a number of reasons, including lens distortion (e.g., barrel distortion, pincushion distortion), rolling shutter correction, geometric transformation (e.g., translational, rotational, Euclidean, affine, projective), motion compensation, electronic image stabilization (EIS), and so forth. The COWEmay process an input image warped by lens distortion into a corrected output image by utilizing a predetermined transformation function that is based on a specific shape and material of a camera lens. For example, a manufacturer of computing devices having a camera may take an average or median measurement of curvature and material makeup of a sample of lenses to predetermine the transformation function. The COWEmay similarly process an input image warped by rolling shutter correction into a corrected output image. The COWEmay process an input image warped by geometric transformation, motion compensation, or EIS by utilizing sensor data (e.g., accelerometer data). The COWEmay determine an arbitrary transformation function based on the sensor data to process the warped input image into a corrected output image.

1 FIG. 1 FIG. 116 102 114 100 116 112 1 112 1 112 1 118 116 102 116 118 further illustrates a userof the computing devicehaving the COWE. In the example environment, the userwishes to capture a photograph of a scene that includes a tree in a foreground and a mountain range in a background. As illustrated by the display-, the scene is rotationally warped in this example and may be referred to as a warped input image. Although rotational warping is illustrated inby the display-, the warped input image may be warped by any number of image distortions, including barrel distortion, pincushion distortion, perspective distortion, skew, curved horizon distortion, panorama distortion, a combination distortion (e.g., a complex distortion), and so forth. A specific image distortion may be caused by a lens of a camera, an unsteady capture of an image or video, and so forth. The display-further illustrates a shutter button, which may be presented to the userby a GUI of the computing device. The usermay tap the shutter buttonto capture an image of the scene.

112 2 114 102 114 116 118 114 104 106 108 The display-illustrates that the warped input image of the scene is corrected by the COWEof the computing deviceinto a corrected output image. The COWEmay correct the warped input image responsive to the usertapping the shutter button. The COWEmay utilize the first cache memory, the second cache memory, and the system memoryto correct the warped input image.

114 116 118 116 118 112 1 114 As an example, the COWEmay receive the warped input image as a result of the usertapping the shutter buttonto capture the scene. The usermay have been unsteady when tapping the shutter button, resulting in the scene being rotationally warped into the warped input image illustrated by the display-. The warped input image may include a plurality of pixel information (e.g., color, location, gamma). The COWEmay divide the warped input image into a number of rows (e.g., two, five, 10, 20) that are as wide as the warped input image and as tall as an integer division of the warped input image. For example, the warped input image may be 1,920 pixels wide and 1,080 pixels tall. In this example, each one of the number of rows is 1,920 pixels wide and an integer division of 1,080 pixels tall. The integer may be two, four, eight, 16, 32, and so forth, resulting in each one of the number of rows being 540 pixels, 270 pixels, 135 pixels, 68 pixels, or 34 pixels, respectively, tall.

114 114 108 After dividing the warped input image into the number of rows, the COWEmay scan, following a coordinate sequence including the number of rows, the plurality of pixel information of the warped input image. The coordinate sequence may include an intra-row coordinate sequence that progresses from bottom to top and left to right within a row and an inter-row coordinate sequence that progresses from a topmost row to bottommost row in a row-by-row fashion. The COWEmay scan the pixel information by accessing it from the system memoryin increments of memory access units (MAUs).

114 108 108 An MAU may be any one of an appropriate dimension, including a number of pixels tall (e.g., four, eight, 16, 32, 60) and another number of pixels wide (e.g., four, six, 10, 12), and may be measured in bits or bytes (B) (e.g., increments of eight bits). A specific dimension of the MAUs accessed by the COWEmay depend on specific information included in the plurality of pixel information of the warped input image and a format of the system memory. The specific information may include color depth (e.g., 8-bit color, 128-bit color, 256-bit color), gamma, and location information. Additionally or alternatively, the pixel information may be compressed (e.g., reducing a number of bytes of an MAU) or uncompressed (e.g., increasing a number of bytes of an MAU). The format of the system memorymay include a bus width (e.g., eight bits, 16 bits, 32 bits) and a clock speed.

As a specific example, using a frame bandwidth reduction (FBR) format, an MAU may be 256 B and include 256 8-bit (e.g., 1 B) pixels. In another example, not using the FBR format, an MAU may still be 256 B but include just 32 64-bit (e.g., 8 B) pixels. As yet another example, a pixel may be described by 30 bits within a 32-bit container and an associated MAU may include 48 pixels, totaling 192 B of pixel information.

114 104 106 114 104 106 108 106 108 104 108 Following the coordinate sequence mentioned above, the COWEmay load portions of the plurality of pixel information of the warped input image into the first cache memory, the second cache memory, or both. The COWEmay do so because the portions of the plurality of pixel information may be referenced multiple times to process pixels of the warped input image into corrected pixels of the corrected output image. Additionally, accessing pixel information from the first cache memoryor the second cache memorymay be significantly faster than accessing pixel information from the system memory. For example, accessing pixel information from the second cache memorymay be 25 times faster than accessing pixel information from the system memory. Similarly, accessing pixel information from the first cache memorymay be 100 times faster than accessing pixel information from the system memory.

114 104 106 114 104 106 108 108 104 Specifically, the COWEmay load a first portion of intra-row pixel information into the first cache memoryand a second portion of inter-row pixel information into the second cache memory. The first portion of pixel information may include a number of bottom-to-top scan lines (e.g., according to the coordinate sequence mentioned above) consisting of one or more MAUs, depending on a height of a row. The second portion of pixel information may include a number of MAUs that form a border between a top row and an adjacent row below the top row. The first portion of pixel information may be referenced most frequently by the COWEto determine corrected output pixels, thus justifying why the first portion may be stored in the first cache memory, which is faster than the second cache memoryand the system memory. Similarly, the second portion of pixel information may be referenced more frequently than portions in the system memorybut less frequently than the first portion in the first cache memory.

114 104 112 2 114 114 104 The COWEmay determine, based on the first portion of pixel information in the first cache memory, a first corrected portion of a plurality of pixel information of the corrected output image illustrated by the display-. The COWEmay do so by determining a transformation function from a warped input pixel location to a corresponding corrected output pixel location. The transformation function may include a predetermined function based on a curvature of a lens of a camera that captured the warped input image. Additionally or alternatively, the transformation function may be based on sensor data (e.g., accelerometer) that is captured with the warped input image. The COWEmay then determine the first corrected portion (e.g., color, gamma) of the plurality of pixel information for the corrected output pixel location using interpolation methods (e.g., linear interpolation, bilinear interpolation, cubic interpolation) on the first portion of pixel information in the first cache memory.

114 106 114 114 114 108 102 The COWEmay determine, based on the second portion of pixel information in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image. The COWEmay similarly do so by determining a transformation function from a warped input pixel location in the second portion of pixel information to a corresponding corrected output pixel location. By performing the method described, the COWEmay determine remaining corrected portions of the plurality of pixel information of the corrected output image. The COWEmay then store the corrected output image in the system memoryof the computing device.

104 106 102 102 116 By following the coordinate sequence mentioned above, sizes of the first cache memoryand the second cache memorymay be optimized. Accordingly, the computing devicemay benefit from power savings and size reduction while correcting warped input images into corrected output images. These benefits enable the computing deviceto be physically smaller and consume less energy, improving a user experience for the user.

2 FIG. 1 FIG. 2 FIG. 200 102 114 102 102 202 202 202 202 202 202 202 202 202 102 102 102 102 a b c d e f g h i In more detail,illustrates an example implementationof the computing devicefrom, which is configured to provide the COWE. The computing deviceis illustrated as various example devices. As non-limiting examples, the computing devicecan be a smartphone, a tablet, a laptop, a desktop, a smartwatch, a pair of smart glasses, a game controller, a smart home speaker, or a microwave appliance. Although not illustrated, the computing devicemay also be implemented as a health monitoring device, a personal media device, a drone, a home appliance, a security system or device thereof, a digital photo frame, and so forth. The computing devicecan be wearable, non-wearable but mobile, or relatively immobile. Further, the computing devicecan be used with or embedded within many computing devices or peripherals (e.g., vehicles, personal computers). The computing devicemay also include additional interfaces or components omitted from.

2 FIG. 1 FIG. 2 FIG. 102 104 106 108 110 112 114 102 202 202 204 206 204 206 202 208 208 210 202 110 202 108 202 204 206 illustrates that the computing deviceincludes various components described with reference to, including the first cache memory, the second cache memory, the system memory, the processor, the display, and the COWE.further illustrates that the computing devicemay include computer-readable media(CRM), which may include memory mediaand storage media. The memory mediamay include one or more non-transitory storage devices, including RAM or DRAM. The storage mediamay include one or more transitory storage devices, including an SSD or a magnetic spinning HDD. The CRMmay further include an operating system(OS) and applications, which may be stored as computer-readable instructions on the CRM. The processorcan execute the computer-readable instructions on the CRMto provide some or all of the functionalities described herein. Although shown as a separate component, the system memorymay be included in the CRMas a standalone memory component, part of the memory media, part of the storage media, or part of both.

2 FIG. 102 212 212 212 114 114 212 102 202 a also illustrates that the computing deviceincludes one or more sensors. The sensorscan include image sensors, audio sensors (e.g., microphones), accelerometers, barometers, ambient light sensors, thermometers, and so forth. One or more of the sensors(e.g., an accelerometer) may be utilized by the COWEto determine transformation functions. Additionally or alternatively, the COWEmay receive a warped input image from one or more of the sensors(e.g., an image sensor, a camera). For example, the computing devicemay be the smartphone, which may include a camera or camera system (e.g., image sensor and lens).

114 114 102 114 102 1 FIG. 2 FIG. In implementations, the COWEcan include one or more integrated circuits (ICs) (e.g., a power management integrated circuit (PMIC)), a system-on-a-chip (SOC), a secure key store, hardware embedded with firmware, a PCB with various hardware components (e.g., a motherboard, a daughterboard), or any combination thereof. As described herein, the COWEmay include one or more components of the computing device, as illustrated inand, configured to determine a corrected output image based on a warped input image. In other implementations, the COWEmay be implemented as the computing device.

102 102 102 Although not shown, the computing devicecan also include input/output (I/O) ports, a system bus, an interconnect, or another data transfer system that couples with various components of or within the computing device. As an example, the I/O ports can enable the computing deviceto interact with other devices or users through peripheral devices, transmitting any combination of digital signals and/or analog signals via wired manners (e.g., ethernet) or wireless manners (e.g., radio). The I/O ports may include any combination of internal or external ports, including universal serial bus ports, audio ports, video ports, and so forth. Various peripheral devices (e.g., human input devices, external CRM, speakers, displays) may be coupled with the I/O ports.

3 FIG. 1 FIG. 2 FIG. 1 2 FIGS.and 1 2 FIGS.and/or 1 2 FIGS.and/or 1 2 FIGS.and/or 2 FIG. 1 2 FIGS.and/or 300 114 300 114 300 104 106 108 202 300 102 depicts an example methodthat a COWE (e.g., COWEofand/or) may implement. The example methodis illustrated as a flowchart with various shapes representing method steps or data (e.g., a warped input image, pixel information). The COWE is like the COWEfrom, except as detailed below. Accordingly, the COWE implementing the example methodmay be operably coupled to a first cache memory (e.g., first cache memoryof), a second cache memory (e.g., second cache memoryof), and/or system memory (e.g., system memoryof, CRMof). Alternatively or additionally, the COWE implementing the example methodmay either be a component of or realized as a computing device (e.g., computing deviceof).

300 302 304 302 102 302 116 3 FIG. 1 FIG. As illustrated, the example methodofincludes a transformation modeland a distortion grid buffer. The transformation modelmay be based on a predetermined set of transformations associated with one or more camera lenses of a computing device (e.g., computing device). For example, a COWE of a computing device having a telephoto lens and a macrophotography lens may include a transformation associated with the telephoto lens and a transformation associated with the macrophotography lens. The transformations may be predetermined based on curvatures of the lenses and configured to correct associated lens distortions. Alternatively or additionally, the transformation modelmay be based on accelerometer data. For example, a user (e.g., userof) of a computing device having a camera, an accelerometer, and a COWE may take a photo of a scene. The user may be unsteady when capturing the photo, which could result in a skewed (e.g., tilted) or blurry photo. However, the accelerometer may record data usable by the COWE to correct the skewed or blurry photo.

304 302 304 The distortion grid buffermay be based on the transformation modelor a subsampled distortion value on a sparse grid. The sparse grid may include neighbor pixels of a current pixel of a warped input image. The COWE may utilize the distortion grid bufferto map a location (e.g., x- and y-coordinates) of the current pixel of the warped input image to a location of a pixel of a corrected output image.

300 306 302 304 306 310 316 306 308 The COWE performing the example methodmay generate a transformation using a transformation generator. As illustrated, the COWE may utilize the transformation modeland the distortion grid bufferas inputs to the transformation generator. Again, the generated transformation may be based on a predetermined transformation associated with a curvature of a camera lens or accelerometer data. The generated transformation may be usable by the COWE to map a location of a pixel of a warped input imageto a location of a pixel of a corrected output image. The transformation generated by the transformation generatormay be input to an address generatorof the COWE.

308 310 316 308 310 316 312 The address generatormay utilize the generated transformation to determine a mapping between locations of a pixel of the warped input imageand a pixel of the corrected output image. The address generatormay further utilize the generated transformation to determine mappings between neighbor pixels of the pixel of the warped inputimage and neighbor pixels of the pixel of the corrected output image. The generated transformation and the locations, as well as associated pixel information (e.g., color, gamma), may be input to a pixel data interpolation.

300 312 316 316 300 316 102 1 2 FIGS.and/or The COWE performing the example methodmay utilize the pixel data interpolationto interpolate pixel information for pixels of the corrected output image. For example, the COWE may utilize any one of a variety of appropriate interpolation methods to determine color and gamma information for the pixels of the corrected output image. Some interpolation methods include piecewise constant interpolation, linear interpolation, bilinear interpolation, polynomial interpolation, spline interpolation, mimetic interpolation, cubic interpolation, bicubic interpolation, and so forth. A specific interpolation method may be selected by the COWE performing the methodbased on a desired quality of the corrected output imageor computing capabilities of a computing device (e.g., computing deviceof) having the COWE.

3 FIG. 300 312 314 314 316 316 316 314 316 314 316 As illustrated in, the COWE performing the example methodmay provide the pixel information from the pixel data interpolationas input data to output cropping. At the output cropping, the COWE may crop the corrected output imageby trimming portions of the corrected output image. The COWE may trim upper portions, lower portions, left portions, or right portions of the corrected output imageat the output cropping. The COWE may determine which portions of the corrected output imageto crop at the output croppingbased on a desired aspect ratio, data size, data compression method, and so forth of the corrected output image.

4 FIG. 3 FIG. 1 FIG. 1 2 FIGS.and/or 400 316 114 400 400 illustrates an example of a corrected output image(e.g., corrected output imageof) that may be divided into rows (e.g., as described with reference to) by a COWE (e.g., the COWEof). As illustrated, the corrected output imagemay be W pixels wide and H pixels tall. W and H may be any positive integer and depend on a desired format, data size, compression method, and so forth of the corrected output image. For example, W may be 1,920 pixels and H may be 1,080 pixels (e.g., a 1080p image). As another example, W may be 3,840 pixels and H may be 2,160 pixels (e.g., a 4k image).

4 FIG. 4 FIG. 316 402 402 316 400 further illustrates that the corrected output imageis divided into rows. The rowsare as wide as the image (e.g., W pixels wide) and m pixels tall, where m is less than H and an integer division of H. For example, H may be 1,080 pixels tall and m may be 108 pixels tall (e.g., an integer division of 10). As another example, H may be 2,160 pixels tall and m may be 216 pixels tall (e.g., an integer division of 10). As an additional example, H may be a same 2,160 pixels tall and m may be 108 pixels tall (e.g., an integer division of 20). A specific integer division may be based on a desired configuration for a COWE that determines the corrected output imageor computing capabilities of a computing device having the COWE. Illustrated in, the corrected output imageis divided into N rows, where N is a positive integer and equal to H divided by m. For example, H may be 1,080 pixels and m may be 108 pixels, making N equal to 10.

5 FIG. 3 FIG. 3 4 FIGS.and/or 5 FIG. 5 FIG. 500 310 316 500 502 502 502 502 1 502 2 502 3 502 4 502 500 500 502 illustrates an example of a warped input image(e.g., warped input imageof) that a COWE may process into a corrected output image (e.g., corrected output imageof).illustrates the warped input imageas four warped rows. Although four warped rowsare described, a number of rows can be any positive integer. The warped rowsare illustrated as a first warped row-, a second warped row-, a third warped row-, and a fourth warped row-. As illustrated, the warped rowsare warped in a barrel distortion manner (e.g., as a result of lens distortion). Although barrel lens distortion is illustrated, distortion of the warped input imagecan be any distortion resulting from lens shapes or an instability during capture of the warped input image. Furthermore, although gaps are illustrated between each of the warped rows, the gaps should not be construed to represent anything beyond an arbitrary warping (e.g., barrel distortion) of an input image. Alternatively or additionally, the gaps may not be present in a warped input image or may be included inand following figures for the purposes of brevity, clarity, or explanation of COWEs.

6 FIG. 1 2 FIGS.and/or 6 FIG. 5 FIG. 6 FIG. 6 FIG. 1 2 FIGS.and/or 600 108 502 3 502 4 602 502 1 502 2 102 602 602 602 illustrates, atgenerally, an example of a COWE loading pixel information from a system memory (e.g., system memoryof).includes the third warped row-and the fourth warped row-fromoverlaid onto a gridof MAUs, illustrated inas rectangles having either solid white backgrounds or shaded backgrounds. The first warped row-and the second warped row-are omitted fromfor clarity. The MAUs include pixel information stored in a system memory of a computing device (e.g., computing deviceof). An MAU of the gridof the MAUs may be a number of pixels tall and a number of pixels wide. The pixels may include a number of bits (e.g., eight bits) that may represent color information or gamma information of the pixels. The gridof MAUs may be considered an overlay of an image sensor of a camera, for example, that stores pixel information associated with pixels of an input image. Accordingly, positions of the MAUs within the gridare not significant beyond that the positions are associated with pixel information of the pixels in same positions of the input image.

310 500 316 400 604 604 606 606 606 602 604 606 606 3 FIG. 5 FIG. 3 FIG. 4 FIG. The COWE may correct a warped input image (e.g., warped input imageof, warped input imageof) into a corrected output image (e.g., corrected output imageof, corrected output imageof) by fetching MAUs from the system memory. The COWE may only fetch MAUs that include pixel information of the warped input image. As illustrated, an MAU that is not fetched by the COWE is a non-fetched MAUand represented by a solid white background. Non-fetched MAUs may not contain pixel information associated with pixels of a warped input image. For example, a lens of a camera may focus light onto an image sensor in a circular fashion to capture pixel information of the warped input image. The lens may focus the light onto areas of the image sensor that are not associated with the non-fetched MAUs. Alternatively, an MAU that is fetched by the COWE is a fetched MAUand represented by a shaded background. Continuing with the present example, the fetched MAUsmay contain pixel information associated with the pixels of the warped input image. That is, the lens of the camera may focus the light onto areas of the image sensor that are associated with the fetched MAUs. Thus, the gridof MAUs includes multiple non-fetched MAUsand multiple fetched MAUs. The pixel information included in the fetched MAUsmay be used by the COWE to interpolate pixel information for pixels of a corrected output image.

7 FIG. 6 FIG. 6 FIG. 1 2 FIGS.and/or 7 FIG. 5 FIG. 3 FIG. 4 FIG. 700 604 606 108 500 316 400 illustrates, atgenerally, various examples of MAUs (e.g., non-fetched MAUsof, fetched MAUsof) that a COWE may fetch from system memory (e.g., system memoryof). In, each of the MAUs is bordered by a number at a top of the MAU and a number on a left side of the MAU. These numbers represent pixel counts that define a shape (e.g., a square, a rectangle) and size, in pixels, of the MAU. A data size of the MAU, in bits (b) or bytes (B) (e.g., eight bits), may depend on specific pixel information for the pixels of the MAU and/or a format of the system memory. Data, for example, may include a color depth (e.g., 8-bit color, 16-bit color) and gamma (e.g., brightness) information, depending on a desired quality of an image (e.g., warped input imageof, corrected output imageof, corrected output imageof). A total data size of the MAU is equal to a number of the pixels in the MAU (e.g., height times width, in pixels) multiplied by a data size of each pixel.

702 702 704 704 704 706 708 710 712 710 714 712 716 714 A first MAUof the various examples of MAUs is a square MAU that is 16 pixels wide and 16 bits tall for a total of 256 pixels. As an example, if the pixel information uses 8-bit color, then the first MAUincludes 2,048 bits (e.g., 256 B) of data. A second MAUis a rectangular MAU that is 32 pixels wide and 8 pixels tall for a total of 256 pixels. As an example, if the pixel information is described using 8 bits, then the second MAUincludes 1,024 bits (e.g., 128 B) of data. The second MAUthat is 32 pixels wide and 8 pixels tall may be utilized in frame rate control (FRC) applications. A third MAUis 16 pixels wide and 8 pixels tall for a total of 128 pixels. A fourth MAUis 16 pixels wide and 4 pixels tall for a total of 74 pixels. A fifth MAUis a rectangular MAU that is 64 pixels wide and one pixel tall. A sixth MAUis half the size of the fifth MAU, being 32 pixels wide and one pixel tall. A seventh MAUis a same size as the sixth MAUbut only 16 pixels wide and two pixels tall. An eighth MAUis half the size of the seventh MAUat 16 pixels wide and only one pixel tall.

8 FIG. 7 FIG. 1 2 FIGS.and/or 1 2 FIGS.and/or 8 FIG. 5 FIG. 6 FIG. 1 FIG. 800 702 716 114 104 106 502 3 502 4 802 804 806 808 604 illustrates, atgenerally, an example of MAUs (e.g., any one of MAUsthroughof) that a COWE (e.g., COWEof) loads into a cache memory (e.g., first cache memoryand/or second cache memoryof).includes the third warped row-and the fourth warped row-fromoverlaid onto a gridof MAUs. As illustrated, there are three categories of MAUs, including inter-row MAUs, intra-row MAUs, and non-fetched MAUs(e.g., non-fetched MAUsof). The COWE may fetch MAUs by following a coordinate sequence (e.g., the coordinate sequence described with reference to).

502 3 500 810 5 FIG. 5 FIG. 8 FIG. The coordinate sequence may include an intra-row coordinate sequence that progresses from bottom to top and left to right or right to left within a row and an inter-row coordinate sequence that progresses from a topmost row to a bottommost row in a row-by-row fashion. Said differently, pixel information of the topmost row may be scanned following an intra-row coordinate sequence that is from bottom to top and left to right or right to left within the topmost row. Then, pixel information of a row adjacent and below the topmost row may be scanned using a same intra-row coordinate sequence. Pixel information of subsequent rows may be similarly scanned. Alternatively, the coordinate sequence may include an intra-row coordinate sequence that progresses from top to bottom and left to right or right to left within a row and an inter-row coordinate sequence that progresses from a bottommost row to a topmost row in a row-by-row fashion. Note that because warped rows (e.g., the third warped row-of) are as wide as warped input images (e.g., warped input imageof), an inter-row coordinate sequence may progress from top to bottom or bottom to top and not left to right or right to left. For clarity and brevity,illustrates a coordinate sequencethat includes an intra-row coordinate sequence from bottom to top and left to right and an inter-row coordinate sequence that is top to bottom.

502 3 502 4 502 3 502 3 810 1 804 1 806 1 806 2 806 3 804 806 1 FIG. The COWE may scan pixels of the third warped row-and the fourth warped row-and fetch associated MAUs from system memory. In this example, the COWE scans pixels of the third warped row-from bottom to top, starting at the left of the third warped row-, along a scan line-(e.g., a bottom to top portion of the coordinate sequence described with reference to). The COWE fetches, from system memory, associated MAUs, including a first inter-row MAU-, a first intra-row MAU-, a second intra-row MAU-, and a third intra-row MAU-. As illustrated, inter-row MAUsare indicated by a vertically striped shading and intra-row MAUsare indicated by a light shading.

804 1 806 1 806 2 806 3 804 1 106 102 806 1 806 2 806 3 104 804 1 804 2 806 1 2 FIGS.and/or 1 2 FIGS.and/or 1 2 FIGS.and/or The COWE may fetch the first inter-row MAU-, the first intra-row MAU-, the second intra-row MAU-, and the third intra-row MAU-using a cache-able read command. For example, the COWE may fetch and load the first inter-row MAU-into a second cache memory (e.g., second cache memoryof, an L2 cache) of a computing device (e.g., computing deviceof). Similarly, the COWE may fetch and load the first intra-row MAU-, the second intra-row MAU-, and the third intra-row MAU-into a first cache memory (e.g., first cache memoryof, an L1 cache) of the computing device. The COWE may repeat this process where a first MAU (e.g., bottom-most MAU, first inter-row MAU-, second inter-row MAU-) of a scan line is loaded into the second cache memory and subsequent MAUs (e.g., intra-row MAUs) of the scan line are loaded into the first cache memory. By so doing, the COWE may increase a speed at which a warped input image is processed into a corrected output image as detailed below.

802 810 810 1 As an example, the MAUs of the gridmay be eight pixels tall by 32 pixels wide for a total of 256 pixels. Further, the COWE may utilize bicubic interpolation on warped input pixels to determine a corrected output pixel. That is, the COWE may bicubically interpolate a warped input pixel into a corrected output pixel based on the warped input pixel and neighboring warped input pixels within a two-pixel radius. As the COWE scans pixels along the scan lines, the COWE may check for MAUs containing pixel information first in the first cache memory and then, if the MAU is not present in the first cache memory, in the second cache memory. In the present example, pixel information for the MAUs associated with the first scan line-is loaded into the first cache memory and the second cache memory. Accordingly, the COWE may access pixel information for the bicubic interpolation from either the first cache memory or the second cache memory, significantly decreasing access times for that pixel information.

8 FIG. 812 1 812 2 816 1 810 2 812 1 806 806 1 812 1 806 804 806 As illustrated,further includes a first warped input pixel-and a second warped input pixel-(not to scale). The COWE may load MAUs associated with the first warped input pixel-as the COWE scans pixels along a second scan line-. As illustrated, the first warped input pixel-is at a left edge of an intra-row MAU, meaning that pixel information for neighboring pixels within a two-pixel radius to the left is included in the first intra-row MAU-that is stored in the first cache memory. The COWE may access that pixel information from the first cache memory, which may be significantly faster (e.g., 100×) than accessing same pixel information from system memory. Thus, the COWE may bicubically interpolate the first warped input pixel-into a corrected output pixel using neighboring pixel information more quickly. Because the COWE may access pixel information of the intra-row MAUsmore often than pixel information of the inter-row MAUs, the pixel information of the intra-row MAUsare stored in the first cache memory.

502 3 502 3 502 4 502 3 502 4 Although not shown, the COWE may continue to scan pixels within the third warped row-using the intra-row coordinate sequence that progresses from bottom to top and left to right. The COWE may scan rightmost pixel information of the third warped row-and then progress to scan (e.g., using a same intra-row coordinate sequence) pixel information within the fourth warped row-. In this way, the COWE may use an inter-row coordinate sequence that progresses from top to bottom. Said differently, the COWE may scan pixel information within an upper row (e.g., the third warped row-) using the intra-row coordinate sequence then scan pixel information within an adjacent row below the upper row (e.g., the fourth warped row-) using the intra-row coordinate sequence.

812 2 806 502 4 812 2 812 2 804 1 810 1 812 2 804 2 The second warped input pixel-, as illustrated, is a top-most pixel of an intra-row MAUwithin the fourth warped row-. Accordingly, to process the second warped input pixel-into a corrected output pixel, the COWE must bicubically interpolate pixel information for neighboring pixels within a two-pixel radius. In this example, the COWE may access neighboring pixel information above the second warped input pixel-within the second cache memory. This is because the first inter-row MAU-was loaded previously into the second cache memory when the COWE scanned pixel information along the first scan line-. Thus, the COWE may process the second warped input pixel-into a corrected output pixel more quickly (e.g., 10×) than if the COWE accessed same neighboring pixel information from system memory. Although not shown, the COWE may similarly process another warped input pixel into another corrected output pixel using the pixel information of a second inter-row MAU-that may be stored in the second cache memory.

Further, due to the coordinate sequence, the COWE may optimize the first cache memory and the second cache memory. The optimization of the first cache memory may be described by Equation 1.

size In Equation 1, L1is the size of the first cache memory, r is the height of a corrected output row, s×L is the MAU shape where s is a short side (e.g., height) and L is a long side (e.g., width), and 4 is the bicubic interpolation window size. Similarly, the optimization of the second cache memory may be described by Equation 2.

size 812 In Equation 2, L2is the size of the first cache memory, b is a bottom line of pixels (e.g., bottom-most pixels within inter-row MAUs), 2 is an interpolation kernel extension, h×w is the MAU shape where h is the height (e.g., eight pixels) and w is the width (e.g., 32 pixels), and H×W is a corrected output image shape where H is the height (e.g., 1,080 pixels) and W is the width (e.g., 1,920 pixels).

As an example, the MAU shape may be eight pixels tall by 32 pixels wide, the corrected output image may be 1,080 pixels tall by 1,920 pixels wide, each row dividing the corrected output image may be 108 pixels tall, and each pixel (e.g., color and gamma) may be described by eight bits. Continuing with the present example, the size of the first cache memory may be optimized to 7,680 pixels or 7,680 B (e.g., 7,680 pixels×8 bits per pixel). Similarly, the size of the second cache memory may be optimized to 99,840 pixels or 99,840 B (e.g., 99,840 pixels×8 bits per pixel).

9 FIG. 3 FIG. 1 2 FIGS.and/or 1 2 FIGS.and/or 900 310 902 104 106 depicts an example methodthat a COWE may perform on a warped input image (e.g., warped input imageof). At, a computing device having the COWE, a first cache memory (e.g., first cache memoryof), and a second cache memory (e.g., second cache memoryof) receives a warped input image comprising a plurality of pixel information. The first cache memory may be an L1 cache that is approximately 100 times faster than system memory (e.g., DRAM, double data rate (DDR) DRAM). The second cache memory may be an L2 cache that is approximately 10 times faster than system memory. The warped input image may be distorted from lens (e.g., of a camera) distortion or an instability (e.g., unsteadiness, long exposure time) during capture of the warped input image.

904 At, the COWE divides the warped input image comprising the plurality of pixel information into N rows. N may be a positive integer, including four, eight, 10, 12, 16, 40, and so forth. The plurality of pixel information may be included in an MAU of a specific size, in pixels, depending on computing capabilities of the computing device and/or a format of system memory of the computing device.

906 At, the COWE scans, following a coordinate sequence including the N rows, the plurality of pixel information of the warped input image. The COWE may scan pixels by accessing MAUs associated with the pixels from system memory of the computing device. The COWE may access individual MAUs or multiple MAUs, for example, from the system memory. The coordinate sequence may include an intra-row coordinate sequence that is at least one of bottom to top and left to right or bottom to top and right to left, as well as an inter-row coordinate sequence that is top to bottom. Alternatively, the coordinate sequence may include an intra-row coordinate sequence that is at least one of top to bottom and left to right or top to bottom and right to left, as well as an inter-row coordinate sequence that is bottom to top.

908 At, the COWE loads a first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device. The first portion may include one or more intra-row MAUs that the COWE accesses from the system memory of the computing device.

910 At, the COWE determines, based on the first portion of the plurality of pixel information of the warped input image stored in the first cache memory, a first corrected portion of a plurality of pixel information of a corrected output image. The COWE may determine the first corrected portion by determining a location of the corrected portion of the plurality of pixel information in the corrected output image using a transformation model. The transformation model may be a predetermined transformation function based on, for example, a curvature of a lens of a camera that captures the warped input image. Alternatively or additionally, the transformation model may be based on sensor data (e.g., accelerometer data) captured contemporaneously with the warped input image. The sensor data may be usable by the COWE to determine the first corrected portion via an EIS process.

912 At, the COWE loads a second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device. The second portion may include one or more inter-row MAUs that the COWE accesses from the system memory of the computing device.

914 At, the COWE determines, based on the second portion of the plurality of pixel information of the warped input image stored in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image. The COWE may determine the second corrected portion using a transformation model and/or sensor data as described above.

916 At, the COWE stores the first portion and the second portion of the plurality of pixel information of the corrected output image as the output image in the system memory of the computing device. The first portion and the second portion can include subsequent first portions and second portions of the corrected output image that the COWE determines by following the coordinate sequence to process all portions of the warped input image. The system memory may be a DDR DRAM or SSD of the computing device.

900 By performing the example methoddescribed above, the COWE may correct a warped input image into a corrected output image more quickly using the first cache memory and the second cache memory than if the COWE used only the system memory. Further, the COWE may optimize sizes of the first cache memory and the second cache memory by following the coordinate sequence described above. By so doing, the COWE improves a user experience of the computing device through smaller physical size and higher power efficiencies of the first and second cache memories.

Example 1: A method comprising: receiving, at a computing device having a cache optimized warp engine, a first cache memory, and a second cache memory, a warped input image comprising a plurality of pixel information; dividing, by the cache optimized warp engine, the warped input image comprising the plurality of pixel information into N rows; scanning, by the cache optimized warp engine and following a coordinate sequence including the N rows, the plurality of pixel information of the warped input image, the coordinate sequence comprising at least one of: an intra row coordinate sequence that is at least one of: bottom to top and left to right; or bottom to top and right to left; and an inter row coordinate sequence that is top to bottom; or an intra row coordinate sequence that is at least one of: top to bottom and left to right; or top to bottom and right to left; and an inter row coordinate sequence that is bottom to top; loading, by the cache optimized warp engine, a first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device; determining, by the cache optimized warp engine and based on the first portion of the plurality of pixel information of the warped input image stored in the first cache memory, a first corrected portion of a plurality of pixel information of a corrected output image; loading, by the cache optimized warp engine, a second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device; determining, by the cache optimized warp engine and based on the second portion of the plurality of pixel information of the warped input image stored in the second cache memory, a second corrected portion of the plurality of pixel information of the corrected output image; and storing, by the cache optimized warp engine, the first portion and the second portion of the plurality of pixel information of the corrected output image as the output image in a system memory of or associated with the computing device. Example 2: The method of example 1, wherein: the warped input image is W pixels wide and T pixels tall; W is an integer greater than one; and T is an integer greater than one. Example 3: The method of example 2, wherein: each row of the N rows is W pixels wide; and each row of the N rows is an integer division of T pixels tall. Example 4: The method of any one of the preceding examples, wherein the coordinate sequence comprises an intra row coordinate sequence that is bottom to top and left to right and an inter row coordinate sequence that is top to bottom. Example 5: The method of any one of the preceding examples, wherein: the first cache is an L1 cache of the cache optimized warp engine; and the second cache is an L2 cache of the cache optimized warp engine; or the second cache is a system-level cache of a system-on-a-chip of the computing device. Example 6: The method of any one of the preceding examples, wherein at least one of: determining the first portion of the plurality of pixel information of the corrected output image uses an arbitrary transform function; or determining the second portion of the plurality of pixel information of the corrected output image uses the arbitrary transform function. Example 7: The method of any one of the preceding examples, wherein determining the first or second portion of the plurality of pixel information of the corrected output image uses at least one of: piecewise interpolation; linear interpolation; polynomial interpolation; spline interpolation; or mimetic interpolation. Example 8: The method of any one of the preceding examples, wherein determining the first or second portion of the plurality of pixel information of the corrected output image uses at least one of: lens distortion; geographic distortion; motion compensation; electronic image stabilization; rolling shutter correction; or camera calibration. Example 9: The method of any one of the preceding examples, wherein at least one of: scanning the plurality of pixel information of the warped input image includes scanning a plurality of memory access units; loading the first portion of the plurality of pixel information of the warped input image into the first cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units; or loading the second portion of the plurality of pixel information of the warped input image into the second cache memory of the computing device includes loading at least one memory access unit of the plurality of memory access units. Example 10: The method of example 9, wherein: the at least one memory access unit is L pixels long and S pixels wide; L is an integer greater than one; and S is an integer greater than one. Example 11: The method of example 10, wherein S is less than L. Example 12: The method of any one of the preceding examples, wherein: a pixel of the warped input image comprises M bits; and M is at least one of: an integer greater than or equal to one; four bits; eight bits; 16 bits; 32 bits; or 64 bits. Example 13: The method of any one of the preceding examples, wherein at least one of: the plurality of pixel information of the warped input image comprises: location information; gamma information; or color information; or the plurality of pixel information of the corrected output image comprises: location information; gamma information; or color information. Example 14: A computing device comprising: a first cache memory; a second cache memory; system memory; at least one processor; and computer readable media storing instructions that, when executed by the at least one processor, cause the at least one processor to implement a cache optimized warp engine utilizing the first cache, the second cache, and the system memory by performing the method of any one of examples 1-13. Example 15: Computer readable media comprising instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of examples 1-13. In the following section, additional examples are provided.

Unless context dictates otherwise, use herein of the word “or” may be considered use of an “inclusive or,” or a term that permits inclusion or application of one or more items that are linked by the word “or” (e.g., a phrase “A or B” may be interpreted as permitting just “A,” as permitting just “B,” or as permitting both “A” and “B”). Also, as used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. For instance, “at least one of a, b, or c” can cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c, or any other ordering of a, b, and c). Further, items represented in the accompanying Drawings and terms discussed herein may be indicative of one or more items or terms, and thus reference may be made interchangeably to single or plural forms of the items and terms in this written description.

Although implementations of systems and techniques of, and apparatuses enabling, cache-optimized warp engines have been described in language specific to certain features and/or methods, the subject of the appended Claims is not necessarily limited to the specific features or methods described. Rather, the specific features and methods are disclosed as example implementations of decoding RF signal reflections into object embeddings for contextual triggers.

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

August 17, 2023

Publication Date

July 9, 2026

Inventors

Dong Wang
Chi-Chun Lai

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “Cache-Optimized Warp Engines” (US-20260195846-A1). https://patentable.app/patents/US-20260195846-A1

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.