Methods and apparatus for camera focusing for video passthrough devices. Gaze information from a gaze tracking subsystem, either alone or along with depth information from a depth tracking system, may be leveraged to determine depths at which to focus. Gaze information, or a combination of depth and gaze information, may be used. As an alternative, the user can manually control the focus distance. For example, a manual bifocal method may provide two focus distances (near focus and far focus.
Legal claims defining the scope of protection, as filed with the USPTO.
a display configured to display virtual content to at least one of a first eye or a second eye of a pair of eyes; a gaze tracker; at least one camera configured to capture images of a scene; and determine first focus distance information based on gaze information from the gaze tracker and a depth map; determine second focus distance information based on vergence of a left gaze vector from the first eye and a right gaze vector from the second eye as determined by the gaze tracker, wherein the second focus distance information is determined based on a distance between the pair of eyes and an intersection point, from the vergence, of the first gaze vector and the second gaze vector; determine a focus distance from the first focus distance information and the second focus distance information; and direct the camera to focus at the focus distance. a controller comprising one or more processors configured to: . A device, comprising:
claim 1 model the first focus distance information and the second focus distance information as probability distance functions (PDFs); and determine the focus distance from the PDFs. . The device as recited in, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to:
claim 2 . The device as recited in, wherein the PDF corresponding to the first focus distance information indicates two or more possible distances, and wherein, to determine the focus distance from the two PDFs, the controller is configured to select one of the two or more distances that most closely corresponds to a distance indicated by the PDF corresponding to the second focus distance information.
claim 1 collect first focus distance observations based on the gaze information from the gaze tracker and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracker; apply a linear regression function to the collected first and second focus distance observations to generate calibrated vergence distances; and determine the focus distance from the calibrated vergence distances. . The device as recited in, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to:
claim 4 . The device as recited in, wherein said collect first focus distance observations, said collect second focus distance observations, and said apply a linear regression function are performed during an enrollment process for the device.
claim 4 . The device as recited in, wherein the observations are based on real objects in the scene that are imaged by the camera and displayed by the display or virtual objects rendered by the controller and displayed by the display.
claim 1 collect first focus distance observations based on the gaze information from the gaze tracker and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracker; train a model based on the collected first and second focus distance observations; and input at least one focus distance observation to the model, wherein the focus distance is output by the model in response to the input. . The device as recited in, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to:
claim 7 . The device as recited in, wherein said collect first focus distance observations, said collect second focus distance observations, and said train a model are performed during an enrollment process for the device.
claim 1 cause display of one or more targets at known ground truth positions; direct a user to fixate on at least one of the one or more targets; record left and right gaze vectors for the user while fixated on the targets; compute intersection points of the left and right gaze vectors and compare the intersection points with the ground truth positions of respective targets; store results of the comparison as calibrated vergence distances; and determine the focus distance from the calibrated vergence distances. . The device as recited in, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to:
claim 1 . The device as recited in, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system.
determining first focus distance information based on gaze information from a gaze tracker and a depth map; determining second focus distance information based on vergence of a left gaze vector from a first eye of a pair of eyes and a right gaze vector from a second eye of the pair of eyes as determined by the gaze tracker, wherein the second focus distance information includes a distance between the pair of eyes and an intersection point, from the vergence, between the first gaze vector and the second gaze vector; determining a focus distance from the first focus distance information and the second focus distance information including the distance between the pair of eyes and the intersection point between the first gaze vector and the second gaze vector; and directing a camera to focus at the focus distance. performing, by a controller comprising one or more processors: . A method, comprising:
claim 11 modeling the first focus distance information and the second focus distance information as probability distance functions (PDFs); and determining the focus distance from the PDFs. . The method as recited in, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises:
claim 12 . The method as recited in, wherein the PDF corresponding to the first focus distance information indicates two or more possible distances, and wherein determining the focus distance from the two PDFs comprises selecting one of the two or more distances that most closely corresponds to a distance indicated by the PDF corresponding to the second focus distance information.
claim 11 collecting first focus distance observations based on the gaze information from the gaze tracker and the depth map; collecting second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracker; applying a linear regression function to the collected first and second focus distance observations to generate calibrated vergence distances; and determining the focus distance from the calibrated vergence distances. . The method as recited in, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises:
claim 14 . The method as recited in, wherein the observations are based on real objects in a scene that are imaged by the camera and displayed by a display or virtual objects rendered by the controller and displayed by the display.
claim 11 collecting first focus distance observations based on the gaze information from the gaze tracker and the depth map; collecting second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracker; training a model based on the collected first and second focus distance observations; and inputting at least one focus distance observation to the model, wherein the focus distance is output by the model in response to the input. . The method as recited in, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises:
claim 16 . The method as recited in, wherein said collect first focus distance observations, said collect second focus distance observations, and said train a model are performed during an enrollment process for the device.
claim 11 displaying one or more targets at known ground truth positions; directing a user to fixate on at least one of the one or more targets; recording left and right gaze vectors for the user while fixated on the targets; computing intersection points of the left and right gaze vectors and compare the intersection points with the ground truth positions of respective targets; storing results of the comparison as calibrated vergence distances; and determining the focus distance from the calibrated vergence distances. . The method as recited in, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises:
claim 11 . The method as recited in, wherein the controller, a display, the gaze tracker, and the camera are components of a head-mounted device (HMD) of an extended reality (XR) system.
a display configured to display virtual content; a gaze tracker; at least one camera configured to capture images of a scene; and determine first focus distance information based on gaze information from the gaze tracker and a depth map; determine second focus distance information based on vergence of a left gaze vector from a first eye of a pair of eyes and a right gaze vector from a second eye of the pair of eyes as determined by the gaze tracker, wherein the second focus distance information includes a distance between the pair of eyes and an intersection point, from the vergence, between the first gaze vector and the second gaze vector; determine a focus distance from the first focus distance information and the second focus distance information including the distance between the pair of eyes and the intersection point between the first gaze vector and the second gaze vector; and direct the camera to focus at the focus distance. a controller comprising one or more processors configured to: a head-mounted device (HMD), comprising: . A system, comprising:
Complete technical specification and implementation details from the patent document.
This application claims benefit of priority to U.S. Provisional Application Ser. No. 63/585,183, entitled “Camera Focusing for Video Passthrough Systems,” filed Sep. 25, 2023, and which is hereby incorporated herein by reference in its entirety.
Extended reality (XR) systems such as mixed reality (MR) or augmented reality (AR) systems combine computer generated information (referred to as virtual content) with real world images or a real-world view to augment, or add content to, a user's view of the world. XR systems may thus be utilized to provide an interactive user experience for multiple applications, such as applications that add virtual content to a real-time view of the viewer's environment, interacting with virtual training environments, gaming, remotely controlling drones or other mechanical systems, viewing digital media content, interacting with the Internet, or the like.
Various embodiments of methods and apparatus for camera focusing for video passthrough on a device, for example head-mounted devices (HMDs) including but not limited to HMDs used in extended reality (XR) applications and systems, are described. HMDs may include wearable devices such as headsets, helmets, goggles, or glasses. An XR system may include an HMD which may include one or more cameras that may be used to capture still images or video frames of the user's environment. The HMD may include lenses positioned in front of the eyes through which the wearer can view the environment. In XR systems, virtual content may be displayed on or projected onto these lenses to make the virtual content visible to the wearer while still being able to view the real environment through the lenses. An XR system in which world-facing cameras are used to capture video of the environment that is then displayed on display screen(s) in front of the user's eyes may be referred to as a video passthrough system.
In an HMD in which world-facing cameras are used to capture video of the environment that is then displayed on display screen(s) in front of the user's eyes (i.e., a video passthrough system), a challenge is to have a good, sharp image for the user at every working distance. In conventional systems, the lenses are fixed focused, which requires the compromise of a limited depth of focus (DoF) at a certain distance from the HMD. For most tasks, objects and surfaces in an environment that are at a sufficient distance are rendered sharply on the display. However, close objects, for example objects within half a meter or less of the HMD, may appear out-of-focus, blurry, to the user when displayed. For example, if the user holds a cellphone in front of the display, the displayed cellphone may be out of focus.
Various embodiments of methods and apparatus for camera focusing for video passthrough devices (e.g., video passthrough HMDs) are described. In some embodiments, rather than using a fixed focus camera with a limited DoF, a variable focus camera may be used along with various gaze-based techniques for determining the depths to focus at to automatically focus at the different depths, including on objects that are close to the HMD. Thus, embodiments overcome the limitation of conventional video passthrough systems in HMDs.
In some embodiments, gaze information from a gaze tracking subsystem or gaze tracker, either alone or along with depth information from a depth tracking system, may be leveraged to determine depths at which to focus. Several embodiments using gaze information, or a combination of depth and gaze information, are described.
For certain users or a certain segment of the population, the gaze-driven focusing techniques may not work well, for example due to the physiology of the users' eyes. As an alternative, embodiments are also described in which the user can manually control the focus distance. For example, a manual bifocal method is described that may provide two focus distances (near focus and far focus), similar to conventional bifocal lenses in glasses.
This specification includes references to “one embodiment” or “an embodiment.” The appearances of the phrases “in one embodiment” or “in an embodiment” do not necessarily refer to the same embodiment. Particular features, structures, or characteristics may be combined in any suitable manner consistent with this disclosure.
“Comprising.” This term is open-ended. As used in the claims, this term does not foreclose additional structure or steps. Consider a claim that recites: “An apparatus comprising one or more processor units . . . .” Such a claim does not foreclose the apparatus from including additional components (e.g., a network interface unit, graphics circuitry, etc.).
“Configured To.” Various units, circuits, or other components may be described or claimed as “configured to” perform a task or tasks. In such contexts, “configured to” is used to connote structure by indicating that the units/circuits/components include structure (e.g., circuitry) that performs those task or tasks during operation. As such, the unit/circuit/component can be said to be configured to perform the task even when the specified unit/circuit/component is not currently operational (e.g., is not on). The units/circuits/components used with the “configured to” language include hardware-for example, circuits, memory storing program instructions executable to implement the operation, etc. Reciting that a unit/circuit/component is “configured to” perform one or more tasks is expressly intended not to invoke 35 U.S.C. § 112, paragraph (f), for that unit/circuit/component. Additionally, “configured to” can include generic structure (e.g., generic circuitry) that is manipulated by software or firmware (e.g., an FPGA or a general-purpose processor executing software) to operate in manner that is capable of performing the task(s) at issue. “Configure to” may also include adapting a manufacturing process (e.g., a semiconductor fabrication facility) to fabricate devices (e.g., integrated circuits) that are adapted to implement or perform one or more tasks.
“First,” “Second,” etc. As used herein, these terms are used as labels for nouns that they precede, and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.). For example, a buffer circuit may be described herein as performing write operations for “first” and “second” values. The terms “first” and “second” do not necessarily imply that the first value must be written before the second value.
“Based On” or “Dependent On.” As used herein, these terms are used to describe one or more factors that affect a determination. These terms do not foreclose additional factors that may affect a determination. That is, a determination may be solely based on those factors or based, at least in part, on those factors. Consider the phrase “determine A based on B.” While in this case, B is a factor that affects the determination of A, such a phrase does not foreclose the determination of A from also being based on C. In other instances, A may be determined based solely on B.
“Or.” When used in the claims, the term “or” is used as an inclusive or and not as an exclusive or. For example, the phrase “at least one of x, y, or z” means any one of x, y, and z, as well as any combination thereof.
Various embodiments of methods and apparatus for camera focusing for video passthrough on a device, for example head-mounted devices (HMDs) including but not limited to HMDs used in extended reality (XR) applications and systems, are described. HMDs may include wearable devices such as headsets, helmets, goggles, or glasses. An XR system may include an HMD which may include one or more cameras that may be used to capture still images or video frames of the user's environment. The HMD may include lenses positioned in front of the eyes through which the wearer can view the environment. In XR systems, virtual content may be displayed on or projected onto these lenses to make the virtual content visible to the wearer while still being able to view the real environment through the lenses. An XR system in which world-facing cameras are used to capture video of the environment that is then displayed on display screen(s) in front of the user's eyes may be referred to as a video passthrough system.
In at least some systems, the HMD may include gaze tracking technology. In an example gaze tracking subsystem or gaze tracker, one or more infrared (IR) light sources emit IR light towards a user's eye. A portion of the IR light is reflected off the eye and captured by an eye tracking camera. Images captured by the eye tracking camera may be input to a glint and pupil detection process, for example implemented by one or more processors of a controller of the HMD. Results of the process are passed to a gaze estimation process, for example implemented by one or more processors of the controller, to estimate the user's current point of gaze. This method of gaze tracking may be referred to as PCCR (Pupil Center Corneal Reflection) tracking. Note that the gaze tracking may be performed for one or for both eyes. Gaze information may be used for several functions of the HMD, for example, a gaze vector may be used to determine in which direction/angle a user is looking in an environment. As another example, in some embodiments, gaze vectors may be determined for both eyes, and an intersection of the two gaze vectors may be used to determine vergence of the eyes, which may indicate at what or where in an environment the user is looking.
In at least some systems, the HMD may include depth tracking technology. In an example depth tracking system, one or more depth cameras and/or other depth sensors may be used to collect depth data that is processed to determine depth (distance from the HMD) of objects and surfaces in the room. Depth information (e.g., depth maps) may be used for several functions of the HMD, for example, depth information may be used in associating virtual content with objects or surfaces in the environment.
In an HMD in which world-facing cameras are used to capture video of the environment that is then displayed on display screen(s) in front of the user's eyes (i.e., a video passthrough system), a challenge is to have a good, sharp image for the user at every working distance. In conventional systems, the lenses are fixed focused, which requires the compromise of a limited depth of focus (DoF) at a certain distance from the HMD. For most tasks, objects and surfaces in an environment that are at a sufficient distance are rendered sharply on the display. However, close objects, for example objects within half a meter or less of the HMD, may appear out-of-focus, blurry, to the user when displayed. For example, if the user holds a cellphone in front of the display, the displayed cellphone may be out of focus.
Various embodiments of methods and apparatus for camera focusing for video passthrough devices (e.g., video passthrough HMDs) are described. In some embodiments, rather than using a fixed focus camera with a limited DoF, a variable focus camera may be used along with various gaze-based techniques for determining the depths to focus at to automatically focus at the different depths, including on objects that are close to the HMD. Thus, embodiments overcome the limitation of conventional video passthrough systems in HMDs.
In some embodiments, gaze information from a gaze tracking subsystem or gaze tracker, either alone or along with depth information from a depth tracking system, may be leveraged to determine depths at which to focus. Several embodiments using gaze information, or a combination of depth and gaze information, are described.
For certain users or a certain segment of the population, the gaze-driven focusing techniques may not work well, for example due to the physiology of the users' eyes. As an alternative, embodiments are also described in which the user can manually control the focus distance. For example, a manual bifocal method is described that may provide two focus distances (near focus and far focus), similar to conventional bifocal lenses in glasses.
1 FIG. 100 102 102 102 102 102 104 102 102 102 102 104 102 104 102 102 graphically illustrates camera focusing for video pass-through systems, according to some embodiments. An environmentin front of an HMD that uses video passthrough technology may include several objectsor surfaces at different distances. ObjectA may be a close object for example within a half meter of the HMD. The other objectsB-D may be at different, farther distances in the environment. Depth tracking technology may be used to determine depth of the various objects, and a depth map may be constructed. Gaze tracking information (and the depth map) may be used to determine a gaze location, and thus an objectthat the user appears to be looking at, in this case objectA. Depth information for objectA may then be used to drive the autofocus camera to focus at a depth corresponding to objectA. If the gaze locationmoves, for example to be clearly on objectD, then depth information for the new location may be used to drive the autofocus camera to focus at the new depth. Note, however, that ambiguities may arise, for example if the gaze locationis at or near the boundary of objectsA andD in this example.
2 2 FIGS.A andB 2 FIG.A 2 FIG.B 200 220 200 210 230 290 202 212 200 202 210 212 202 212 200 202 210 212 illustrate a depth-based focus method, according to some embodiments. In this embodiment, as shown in, gaze tracking technology is used to determine a gaze vectorfor the user, an intersectionof the gaze vectorwith an object or surface in a depth mapis determined, and depth information corresponding to that object or surface is used to drive the camera to focus at focus distance. As shown in, a userlooks towards objectA at depthA. Gaze vectorA is determined, an intersection with objectA in the depth mapis determined, and the camera is driven to focus at depthA. If the user moves their gaze to look at objectB at depthB, gaze vectorB is determined, an intersection with objectB in the depth mapis determined, and the camera is driven to focus at depthB.
290 210 200 202 212 200 210 Note that as the userturns their head, moves about in the environment, or moves an object (for example, moving their hand that holds a cellphone in front of the HMD), the depth tracking technology dynamically updates the depth map, the gaze vectoris updated, and object/depthinformation determined from the gaze vectorand depth mapmay be dynamically updated, which may in turn drive the camera to continuously and dynamically focus at new depths during use.
This depth-based method may work well in most cases. However, ambiguities may arise, for example if the gaze location is at or near the boundary of objects in a scene, which may cause the autofocus functionality to switch rapidly between different depths. In addition, this method depends on the reliability of the depth information, which may in some cases or conditions not be precisely accurate, and thus may result in focusing at the wrong depth.
3 3 FIGS.A andB 2 2 FIGS.A andB 3 FIG.A 3 FIG.B 308 306 320 306 308 330 230 292 302 312 306 308 312 312 202 212 306 308 312 312 illustrate a vergence-based focus method, according to some embodiments. Instead of using a binocular gaze vector as illustrated in, these embodiments use gaze vectors from both eyes to estimate distance based on vergence of the left and right gaze vectors. The depth map may not be used in these embodiments. As shown in, gaze tracking technology is used to determine a left gaze vectorfor the user's left eye and a right gaze vectorfor the user's right eye, and an intersectionof the gaze vectorsandare used to determine a focus distance, which is used to drive the camera to focus at that focus distance. As shown in, a user's eyeslook towards objectA at depthA. Gaze vectorsA andA are determined, an intersection of the two gaze vectors is determined that indicates depthA, and the camera is driven to focus at depthA. If the user moves their gaze to look at objectB at depthB, gaze vectorsB andB are determined, an intersection of the two gaze vectors is determined that indicates depthB, and the camera is driven to focus at depthB.
2 2 FIGS.A andB This vergence-based method may work well in most cases, and does not depend on depth map information as does the depth-based method described in reference to. In addition, this method may not suffer from the ambiguities at or near the boundary of objects as does the depth-based method. However, the accuracy of this method at determining the precise depth of an object at which the user is looking may in some cases not be as high. In other words, the depth determined by the intersection of the left and right gaze vectors may be a bit off from the actual depth of the object, which may result in the object being somewhat out of focus.
Several embodiments that combine depth and vergence information to determine depth to drive autofocus of the camera are described that may overcome the limitations of the depth-based and vergence-based approaches described above. Depth-based focusing has high accuracy in determining depth but may introduce ambiguity as to what the user is trying to focus on. Vergence-based focusing provides a relatively stable signal for where/on what the user is trying to focus, but the focus depth is not always accurate. The embodiments described below combine the two signals to take advantage of the accuracy of depth-based focusing, as well as the ability of vergence to drive the focus to the right place in a scene.
4 FIG. illustrates a vergence-depth fusion method that uses a probabilistic framework to infer the focus distance, according to some embodiments. Since both the gaze information and the depth information may be “noisy”, treating one or both as certainties can lead to incorrect behavior, such as jumping back and forth between near focus and far focus at the edges of objects. In embodiments of a probabilistic framework method, depth is modeled as a distribution of possible answers of where an object is in a scene that a user might be looking at. Vergence is modeled as an uncertainty both in angular space and in distance. By formulating this as a probabilistic problem, these noisy estimates, distributions of possible answers, can be examined together to find a most likely answer.
4 FIG. 400 402 400 402 In, in the top path (and), gaze and depth information may be represented as a probability density function (PDF). Element, the “cloud” (region of uncertainty) represents the fact that the gaze tracking algorithm may indicate that the gaze vector is landing right on edge of the square. There may be as much as several degrees of error. By modeling as an uncertainty cloud, the gaze/depth information covers at least portions of both the foreground object (the square) and the background object (the triangle). Looking at the distribution of depth (histogram), there may be two hypotheses for what object the user is looking at.
4 FIG. 410 412 In, the bottom path (and) represents vergence. Independent of what is happening in the scene, there is uncertainty of where the user's eyes are actually verging. Vergence is modeled as an uncertainty both in angular space and in distance.
420 430 432 432 432 Atand, the two paths are combined to find a solution for the most likely distance (). A Bayesian method (or some other method) may be used to estimate what is the most likely distancethe user is actually looking at. The depth information (top path) indicates that the user is either looking at either the square or the triangle. The vergence information (bottom path) indicates that the user is probably looking somewhere around the square, but exactly where is not known. The vergence information may thus function as a tie-breaker by lending credibility to the hypothesis that the user is looking at the square, and therefore that should be the solution and the distancethat is focused at.
4 FIG. 432 432 Generally stated, in reference to, a formulation is provided where the noisy depth and vergence signals are fused to find the most likely distancethat the user is currently focusing on in the scene, and that distanceis used to drive the focus of the camera. Mathematically, this may be viewed as an estimation problem. There are two channels of information. Each channel has its own error models, and distribution of errors. This method fuses those two channels in an optimal way to improve the estimation of depth.
5 5 FIGS.A andB 500 510 520 illustrate another vergence-depth fusion method that uses online or offline calibration of a vergence--depth model, according to some embodiments. This method collects depth observations (the generally straight lines) and vergence observations (the generally fuzzy or stochastic lines) () and performs a function on the combined observations (e.g., linear regression) to remove bias, to generate calibrated observations.
5 FIG.A In, the vergence observations show that the eyes are verging at some distance, but are tending to jump back and forth between a near and a far object, and perhaps something in between. Thus, these signals are noisy. The depth observations represent ground truth distances for where the objects are, how far away. There is bias in the signals, an error over time, where the vergence of the eyes is being measured at a wrong distance. Note that the objects may be real objects in the environment imaged by the camera, or virtual objects rendered and displayed on the display screen (e.g., a UI element, text, icon, or any virtual content).
500 510 520 520 These observations may be collected over time (), and input to a linear regression optimizer () to obtain a model of how to map the vergence signal, to pull the vergence signal to the actual depth of the objects in a scene. Once the model, is applied, the bias is removed or reduced from the signals. The average vergence distance is much closer to the actual ground truth depth information across different distances in a scene, as shown at.
This method uses vergence as a primary signal for where to set the focus, but the vergence signal is calibrated against the ground truth of real distances in the scene to make the depth indicated by the vergence signal more accurate. Vergence is generally reliable and accurate when it can be trusted. When reliable, with good confidence, those observations may be used to train or calibrate the vergence signal. After calibration, the vergence signal may be both more reliable and accurate.
5 FIG.A This calibration process may be done for each user of the device (e.g., an HMD) either offline (e.g., during an enrollment process) or online (when the user is actively using the device). (The linear regression curve shown inis unique to each user, as vergence varies among users). In an example offline process, every time there is a new user of an HMD with gaze tracking, the user's eyes are “enrolled” to build a model of the eyes in relation to the “real world” and the device, to be used in gaze tracking. Conventionally, an angular enrollment is performed to determine the angle of gaze. By adding objects at different known distances to the enrollment process, vergence can be enrolled in addition to angular gaze.
The HMD may have depth sensing technology; however, in some situations the depth information may not be reliable. In situations where the depth information can be trusted, and where there is vergence information, observations of vergence together with depth can be recorded. Over time, a sufficient number of reliable observations can be recorded, and the linear regression model can be generated. The system may continually adaptively adjust the model over time as additional observations are recorded to ensure that the model correctly reflects what the user's vergence is doing. In other words, during normal use of the device, when high confidence situations are detected for the depth and vergence signals, those observations may be used as calibration points to improve the user's linear regression curve. In some embodiments, the system may include a confidence map for each depth map. The confidence map may be used to determine the confidence of a depth value, and thus objects that are the most reliable to get training data from may be determined.
5 FIG.B illustrates a vergence enrollment and initial calibration method that uses “targets” displayed to the user, according to some embodiments. In some embodiments, a vergence enrollment process may use simple virtual or real targets (e.g., ball-like targets), similar to what may be used in a conventional eye enrollment process. These targets may be placed at different XYZ or radial angle/azimuthal angle/distances. In some embodiments, a target is positioned at a known ground-truth position (XYZ) from a list of different XYZ or radial angle/azimuthal angle/distances. During enrollment, the user is asked to fixate on a target. The left/right gaze vectors are recorded, and the intersection point is computed and compared with the ground truth position. In various embodiments, regression/machine learning or a look-up table may be used to map computed vergence distance to calibrated vergence distance. In some embodiments, after an initial calibration, online calibration may be performed whenever there are reliable depths from real or virtual objects that the user is fixating on.
6 FIG. 6 FIG. 600 610 620 630 640 650 illustrates another vergence-depth fusion method that uses a generative model, according to some embodiments. In this method, a model may be trained to generate a fused camera focus distance sequence based on vergence and depth sequences. A network may be trained to take a sequence of vergence observations and depth observations and generate a fused sequence from the two input sequences. At a high level, this method uses machine learning to solve the problem. A model may be trained on vergence and depth information, and after training camera focus distances can be output for input vergence and depth inputs. As shown in, a vergence sequenceand depth sequenceare input to an encoderfunction that trains a model. A decoderfunction outputs a fused sequence.
7 7 FIGS.A andB 7 FIG.A 7 FIG.B 700 790 710 720 730 740 720 750 760 770 760 770 730 740 740 750 770 illustrate vergence and closed-loop focus hybrid methods, according to some embodiments. In these methods, vergence distance gets close to a target and defines a focus position optimization range. Closed-loop feedback may then be applied between the focus position and an image sharpness metric to refine the result.graphically illustrates a vergence distancefor a user, and a range of focus distancesin which the distance is to be optimized.shows the closed-loop optimization method. A vergence distance is estimatedand input into a focus control algorithm. A focus positionbased on the input fromis used to drive the camerato the focus distance. An imageis captured, and an image sharpness metricis derived from the image. The image sharpness metricis fed back to the focus control algorithm, which adjusts the focus positionas necessary and uses the new focus positionto drive the camerafocus distance. This feedback loop may continue until the image sharpness metricis optimized.
4 6 FIGS.through This method performs automatic focusing without fusing the depth and vergence information as is done in. Vergence is used to determine roughly where focus, and some metric in image space (e.g., an image sharpness metric) can be leveraged to indicate how sharp the object appears. This is performed in a closed loop, making small adjustments to the focus position, to drive the camera to an optimal focus distance. Conventional autofocus may have a region of interest, and attempt to optimize sharpness or contrast in that region. A difference from conventional autofocus in this method is that the method knows roughly the solution based on vergence, and constrains how much the actuator is moved so that the actuator is not driven to a totally different position to focus on a significantly different object distance. The optimization is “primed” with the vergence distance as a starting point for finding the correct focus distance. The “last mile” of focusing is performed in a closed loop, with an image metric used to optimize the focusing within a narrow range.
770 770 Any of several image metrics, or combinations thereof, may be used in the feedback loop. In some embodiments, the feedback metriccould be a conventional autofocus (AF) metric, but limited to optimize within a range determined from the vergence signal.
7 7 FIGS.C andD 7 7 FIGS.A andB 7 FIG.C 7 FIG.D 7 FIG.B 790 782 780 786 788 illustrate a method in which focus pixels are used to determine whether a region the user is fixating on is in focus, and in which camera focusing is adjusted to make the camera in focus depending on a sign of disparity in focus pixel(s), according to some embodiments. This method may, for example, be used to perform the “last mile” focusing described in reference to. Autofocus sensors are mature technology. In this technology, the sensor has a pixel structure, for example left/right sub-pixels, that can detect if a region in the image is in focus or not by detecting whether there is disparity between left/right sub-pixels. This method is similar to a conventional tap-to-focus method used in smartphones and tablets, but the tap is replaced with gaze (the gaze vector) to determine where in the image the camera is to focus.illustrates a userlooking at an image that includes focus pixels with left/right parity disparity detection. As shown in, the intersection of a gaze vectorwith a focus pixel sensor imageis used to determine an autofocus region of interest (ROI). At, a focusing method is applied (for example, the closed-loop method of) to make the ROI in focus based on sub-pixel left/right disparity.
7 7 FIGS.A throughD 2 8 FIGS.A through Embodiments of the methods as illustrated inmay be used with any of the methods described in reference toto do “last mile” focusing once a starting focus position has been determined according to the respective method.
8 FIG. illustrates a manual bifocal focusing method, according to some embodiments. For certain users or a certain segment of the population, the gaze-driven focusing techniques described above may not work well, for example due to the physiology of the users' eyes. As an alternative, embodiments are also described in which the user can manually control the focus distance. For example, a manual bifocal method is described that may provide two focus distances (near focus and far focus), similar to conventional bifocal lenses in glasses.
8 FIG. 2 7 FIGS.A through Embodiments of the method described in reference tomay provide a way to deliver the benefits of having a focusable camera to users who choose not to use continuous automatic focusing as described in reference to, or to users for which the continuous focusing methods do not work, for example due to the physiology of the users' eyes. As an example, some portion of the population may have eyes for which the vergence does not closely correspond to the depth at which they want to focus. Embodiments may provide a manual method to control the focus distance of the camera to suit particular users' needs or preferences; the method may provide two predetermined focus distances (near focus and far focus), rather than the continuous focus distances as previously described. In some embodiments, the user may select between the two focus modes, for example using a control of or on the HMD or by performing a gesture with the hands or eyes to switch modes.
Note that in some embodiments, the method may be configured to automatically switch between the two preconfigured focus distances based on a detected condition, for example detecting the user looking downwards through the lenses of the HMD rather than straight ahead may be used to automatically switch to near focus mode rather than far focus mode, or the detection of an object intersected by the user's gaze vector that is within a certain minimum distance may be used to automatically switch to near focus mode.
8 FIG. 830 820 810 800 810 800 Referring to, curverepresents the conventional fixed-focus technique. Curvecorresponds to focus using a depth-based technique. Curvesandrepresent splitting the focus into two portions, near and far. Curvecorresponds to a far focus mode, in this example set to approximately 0.7 m. Curvecorresponds to a near focus mode, in this example set to approximately 0.4 m. For any object approximately 0.5 m or farther from the camera, the far focus setting is use. For closer objects, the device may toggle the camera to the near focus setting. By carefully choosing the two values, the whole range can be effectively covered while providing satisfactory sharpness.
Using these methods, users may be given control of the focus distance, with no ambiguity as may be present in the other methods described herein, while allowing the user to focus on near objects that may not be feasible using a conventional fixed-focus technique. This method does not suffer from errors in gaze, vergence, and depth maps that may be present in the other methods described herein, and do not depend on the camera sensor to determine focus distance.
8 FIG. However, these methods may place the burden on the user to manually select the focus distance, rather than providing continuous, automatic focusing as in the other methods described herein. If the user does not select the correct setting, sharpness will be poor, for example as illustrated by the region of regression if the far focus setting is used for close objects as illustrated in.
8 FIG. 2 7 FIGS.A through 8 FIG. 8 FIG. Conventional fixed-focused methods may be viewed as single-plane method. The method ofmay be viewed as a dual-plane method. The automatic methods described in reference toare continuous, many-plane methods. While embodiments are described as dual-plane methods as shown in, a multiple fixed-plane method may be used, such as a trifocal method similar to, but with another, intermediate mode.
2 7 FIGS.A through 8 FIG. In some embodiments, an HMD may support one or more of the continuous focus methods as illustrated in reference to, and may also support the bifocal focus method as illustrated in reference to. A control or setting for the HMD may be used to select the particular focus method that the user wants to use. Thus, a user who does not want to, or that cannot successfully, use one of the continuous focus modes may choose to switch the HMD into bifocal focus mode. In addition, there may be certain conditions, such as low light conditions, where one of the continuous, automatic focusing methods does not work well; a user may choose to switch to bifocal focusing in those conditions, even if the continuous focusing method normally works well for them. In some embodiments, the device may automatically switch between a continuous focusing mode and a bifocal mode upon detecting certain such conditions.
9 FIG. 9 FIG. 2 2 FIGS.A-B 900 910 920 930 930 900 is a high-level flowchart of a depth-based focus method, according to some embodiments, according to some embodiments. The method ofcorresponds to the method shown in. As indicated at, a gaze vector may be estimated by a gaze tracking subsystem or gaze tracker. As indicated at, an intersection of the gaze vector with a depth map may be determined. As indicated at, a focus distance may be determined from an object or surface at the intersection of the gaze vector with the depth map. As indicated at, the camera may be directed to focus at the determined focus distance. As indicated by the arrow returning fromto, the method may continue as long as the device is being used.
10 FIG. 10 FIG. 3 3 FIGS.A-B 1000 1010 1020 1030 1030 1000 is a high-level flowchart of a vergence-based focus method, according to some embodiments. The method ofcorresponds to the method shown in. As indicated at, gaze vectors may be determined for both the left and right eyes by a gaze tracking subsystem or gaze tracker. As indicated at, vergence (an intersection of the gaze vectors in 3D space) may be determined. As indicated at, a focus distance may be determined from the vergence information. As indicated at, the camera may be directed to focus at the determined focus distance. As indicated by the arrow returning fromto, the method may continue as long as the device is being used.
11 FIG. 11 FIG. 4 FIG. 1100 1110 1120 1130 1130 1100 is a high-level flowchart of a vergence-depth fusion method that uses a probabilistic framework to infer the focus distance, according to some embodiments. The method ofcorresponds to the method shown in. As indicated at, focus distance(s) may be estimated based on gaze and depth map information, and a focus distance may also be estimated based on vergence. As indicated at, the two estimates of focal distances may be modeled as probability distance functions (PDFs). As indicated at, a most likely focus distance may be determined by examining the two PDFs together. As indicated at, the camera may be directed to focus at the most likely focus distance. As indicated by the arrow returning fromto, the method may continue as long as the device is being used.
12 FIG. 12 FIG. 5 FIG. 1200 1210 1220 1230 1230 1200 1230 1200 1200 1210 is a high-level flowchart of another vergence-depth fusion method that uses online or offline calibration of a vergence--depth model, according to some embodiments. The method ofcorresponds to the method shown in. As indicated at, distance observations based on depth map information and distance observations based on vergence are collected over time. The distance observations may be collected for real objects in the environment that are captured by the camera, and/or for virtual objects that are rendered and displayed on the display screen. As indicated at, a function (e.g., linear regression) may be applied to the collected observations to generate calibrated vergence distances. As indicated at, a focus distance may be determined from the calibrated vergence distances, for example by receiving a vergence estimate and adjusting the estimate to a corresponding calibrated vergence distance. As indicated at, the camera may be directed to focus at the determined focus distance. As indicated by the arrow returning fromto, the method may continue as long as the device is being used. As indicated by the dashed arrow returning fromto, the calibration process may be performed during use of the device, or a previous calibration may be updated with new information during use of the device. In some embodiments, and initial calibration of vergence distances (-) may be performed during an enrollment process for a user of the device.
12 FIG. In some embodiments, as an alternative to the method shown in, a vergence enrollment process may use simple virtual or real targets (e.g., ball-like targets), similar to what may be used in a conventional eye enrollment process. These targets may be placed at different XYZ or radial angle/azimuthal angle/distances. In some embodiments, a target is positioned at a known ground-truth position (XYZ) from a list of different XYZ or radial angle/azimuthal angle/distances. During enrollment, the user is asked to fixate on a target. The left/right gaze vectors are recorded, and the intersection point is computed and compared with the ground truth position. In various embodiments, regression/machine learning or a look-up table may be used to map computed vergence distance to calibrated vergence distance. In some embodiments, after an initial calibration, online calibration may be performed whenever there are reliable depths from real or virtual objects that the user is fixating on.
13 FIG. 13 FIG. 6 FIG. 1300 1310 1320 1330 1330 1320 is a high-level flowchart of another vergence--depth fusion method that uses a generative model, according to some embodiments. The method ofcorresponds to the method shown in. In this method, a model may be trained to generate a fused camera focus distance sequence based on vergence and depth sequences. A network may be trained to take a sequence of vergence observations and depth observations and generate a fused sequence from the two input sequences. As indicated at, distance observations based on depth map information and distance observations based on vergence may be collected. As indicated at, a model may be trained based on the vergence and depth sequences. As indicated at, during use of the device, a most likely focus distance may be determined by inputting depth and/or vergence information into the model. As indicated at, the camera may be directed to focus at the determined focus distance. As indicated by the arrow returning fromto, the method may continue as long as the device is being used. Note that the model may be updated by new observations made during use to thus improve the model. An initial model may be generated during enrollment of a user on the device.
14 FIG. 14 FIG. 7 7 FIGS.A-B 1400 1410 1420 1430 1430 1410 1410 1440 1430 1430 1400 is a high-level flowchart of a vergence and closed-loop focus hybrid method, according to some embodiments. The method ofcorresponds to the method shown in. In this method, vergence distance gets close to a target and defines a focus position optimization range. Closed-loop feedback may then be applied between the focus position and an image sharpness metric to refine the result. As indicated at, an initial focus distance may be determined based on vergence. As indicated at, the camera may be directed to focus at the focus distance. As indicated at, an image may be captured by the camera at the focus distance, and analyzed to determine a value for a sharpness metric. At, in the sharpness metric has been optimized (e.g., is within a specified acceptable range), the closed-loop focusing is done. At, if the metric has not been optimized, the focus distance may be adjusted (at least in part based on the value of the sharpness metric), and the method returns toto capture and analyze another image at the new focus distance. Elements-may be repeated in a “closed loop” until the test at elementis satisfied. As indicated by the arrow returning fromto, the entire method may be continued as long as the device is in use.
7 7 FIGS.C andD 14 FIG. 14 FIG. In some embodiments, focus pixels are used to determine whether a region the user is fixating on is in focus, and in which camera focusing is adjusted to make the camera in focus depending on a sign of disparity in focus pixel(s), as shown in. This method may, for example, be used to perform the “last mile” focusing described in reference to. In this method, the sensor has a pixel structure, for example left/right sub-pixels, that can detect if a region in the image is in focus or not by detecting whether there is disparity between left/right sub-pixels. In this method, the intersection of a gaze vector with a focus pixel sensor image is used to determine an autofocus region of interest (ROI). At A focusing method is applied (for example, the closed-loop method of) to make the ROI in focus based on sub-pixel left/right disparity.
15 FIG. 15 FIG. 8 FIG. 8 FIG. 8 FIG. 1500 1510 1530 is a high-level flowchart of a manual bifocal method, according to some embodiments. The method ofcorresponds to the method shown in. For certain users or a certain segment of the population, the gaze-driven focusing techniques described above may not work well, for example due to the physiology of the users' eyes. As an alternative, in some embodiments, the user can manually control the focus distance. For example, a manual bifocal method may be implemented that may provide two focus distances (near focus and far focus), similar to conventional bifocal lenses in glasses. As indicated at, the system may focus the camera at a first distance (for example far focus as illustrated in). As indicated at, the system may detect that the user has changed (or wants to change) focus to a second distance (for example near focus as illustrated in). The device may then switch the camera focus to the second distance. As indicated at, focus may remain at this distance until detecting that the user has changed (or wants to change) back to the first distance.
In some embodiments, the user may select between the two focus modes or distances, for example using a control of or on the HMD or by performing a gesture with the hands or eyes to switch modes. In some embodiments, the method may be configured to automatically switch between the two preconfigured focus distances based on a detected condition, for example detecting the user looking downwards through the lenses of the HMD rather than straight ahead may be used to automatically switch to near focus mode rather than far focus mode, or the detection of an object intersected by the user's gaze vector that is within a certain minimum distance may be used to automatically switch to near focus mode.
The far focus mode may, for example be set to approximately 0.7 m. The near focus mode may, for example, be set to approximately 0.4 m. For any object approximately 0.5 m or farther from the camera, the far focus setting may be use. For closer objects, the device may toggle the camera to the near focus setting. By carefully choosing the two values, the whole range can be effectively covered while providing satisfactory sharpness.
16 16 FIGS.A throughC 1 15 FIGS.through 16 16 FIGS.A throughC 16 FIG.A 16 16 FIGS.B andC 16 FIG.A 16 FIG.B 1900 1900 1900 1900 1900 1930 1930 1930 illustrate example devices in which the methods ofmay be implemented, according to some embodiments. Note that the HMDsas illustrated inare given by way of example, and are not intended to be limiting. In various embodiments, the shape, size, and other features of an HMDmay differ, and the locations, numbers, types, and other features of the components of an HMDand of the eye imaging system.shows a side view of an example HMD, andshow alternative front views of example HMDs, withshowing device that has one lensthat covers both eyes andshowing a device that has rightA and leftB lenses.
1900 1930 1910 1900 1900 1900 1920 1900 HMDmay include lens(es), mounted in a wearable housing or frame. HMDmay be worn on a user's head (the “wearer”) so that the lens(es) is disposed in front of the wearer's eyes. In some embodiments, an HMDmay implement any of various types of display technologies or display systems. For example, HMDmay include a display system that directs light that forms images (virtual content) through one or more layers of waveguides in the lens(es); output couplers of the waveguides (e.g., relief gratings or volume holography) may output the light towards the wearer to form images at or near the wearer's eyes. As another example, HMDmay include a direct retinal projector system that directs light towards reflective components of the lens(es); the reflective lens(es) is configured to redirect the light to form images at the wearer's eyes.
1900 1920 1950 1920 1950 1910 1900 1980 In some embodiments, HMDmay also include one or more sensors that collect information about the wearer's environment (video, depth information, lighting information, etc.) and about the wearer (e.g., eye or gaze tracking sensors). The sensors may include one or more of, but are not limited to one or more eye tracking cameras(e.g., infrared (IR) cameras) that capture views of the user's eyes, one or more world-facing or PoV cameras(e.g., RGB video cameras) that can capture images or video of the real-world environment in a field of view in front of the user, and one or more ambient light sensors that capture lighting information for the environment. Camerasandmay be integrated in or attached to the frame. HMDmay also include one or more light sourcessuch as LED or infrared point light sources that emit light (e.g., light in the IR portion of the spectrum) towards the user's eye or eyes.
1960 1900 1900 1960 1960 A controllerfor the XR system may be implemented in the HMD, or alternatively may be implemented at least in part by an external device (e.g., a computing system or handheld device) that is communicatively coupled to HMDvia a wired or wireless interface. Controllermay include one or more of various types of processors, image signal processors (ISPs), graphics processing units (GPUs), coder/decoders (codecs), system on a chip (SOC), CPUs, and/or other components for processing and rendering video and/or images. In some embodiments, controllermay render frames (each frame including a left and right image) that include virtual content based at least in part on inputs obtained from the sensors and from an eye tracking system, and may provide the frames to the display system.
1970 1900 1900 1970 1950 1910 1970 Memoryfor the XR system may be implemented in the HMD, or alternatively may be implemented at least in part by an external device (e.g., a computing system) that is communicatively coupled to HMDvia a wired or wireless interface. The memorymay, for example, be used to record video or images captured by the one or more camerasintegrated in or attached to frame. Memorymay include any type of memory, such as dynamic random-access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2, DDR3, etc.) SDRAM (including mobile versions of the SDRAMs such as mDDR3, etc., or low power versions of the SDRAMs such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. In some embodiments, one or more memory devices may be coupled onto a circuit board to form memory modules such as single inline memory modules (SIMMs), dual inline memory modules (DIMMs), etc. Alternatively, the devices may be mounted with an integrated circuit implementing system in a chip-on-chip configuration, a package-on-package configuration, or a multi-chip module configuration. In some embodiments DRAM may be used as temporary storage of images or video for processing, but other storage options may be used in an HMD to store processed data, such as Flash or other “hard drive” technologies. This other storage may be separate from the externally coupled storage mentioned below.
16 16 FIGS.A throughC 1980 1920 1950 1980 1920 1950 1980 1920 1950 Whileonly show light sourcesand camerasandfor one eye, embodiments may include light sourcesand camerasandfor each eye, and gaze tracking may be performed for both eyes. In addition, the light sources,, eye tracking cameraand PoV cameramay be located elsewhere than shown.
1900 1900 1900 1960 1950 1950 1960 1900 1920 1960 1900 16 16 FIGS.A-C 1 15 FIGS.through Embodiments of an HMDas illustrated inmay, for example, be used in augmented or mixed (AR) applications to provide augmented or mixed reality views to the wearer. HMDmay include one or more sensors, for example located on external surfaces of the HMD, that collect information about the wearer's external environment (video, depth information, lighting information, etc.); the sensors may provide the collected information to controllerof the XR system. The sensors may include one or more visible light cameras(e.g., RGB video cameras) that capture video of the wearer's environment that, in some embodiments, may be used to provide the wearer with a virtual view of their real environment. In some embodiments, video streams of the real environment captured by the visible light camerasmay be processed by the controllerof the HMDto render augmented or mixed reality frames that include virtual content overlaid on the view of the real environment, and the rendered frames may be provided to the display system. In some embodiments, input from the eye tracking cameramay be used in a PCCR gaze tracking process executed by the controllerto track the gaze/pose of the user's eyes for use in rendering the augmented or mixed reality content for display. In addition, one or more of the methods as illustrated inmay be implemented in the HMD to provide camera focusing in video passthrough mode for the HMD.
17 FIG. 1 15 FIGS.through is a block diagram illustrating an example device that may include components and implement methods as illustrated in, according to some embodiments.
2000 2000 2000 2060 2060 In some embodiments, an XR system may include a devicesuch as a headset, helmet, goggles, or glasses. Devicemay implement any of various types of display technologies. For example, devicemay include a transparent or translucent display(e.g., eyeglass lenses) through which the user may view the real environment and a medium integrated with displaythrough which light representative of virtual images is directed to the wearer's eyes to provide an augmented view of reality to the wearer.
2000 2060 2030 2000 2070 2074 2060 2078 2060 2070 2050 2000 2060 In some embodiments, devicemay include a controllerconfigured to implement functionality of the XR system and to generate frames (each frame including a left and right image) that are provided to display. In some embodiments, devicemay also include memoryconfigured to store software (code) of the XR system that is executable by the controller, as well as datathat may be used by the XR system when executing on the controller. In some embodiments, memorymay also be used to store video captured by camera. In some embodiments, devicemay also include one or more interfaces (e.g., a Bluetooth technology interface, USB interface, etc.) configured to communicate with an external device (not shown) via a wired or wireless connection. In some embodiments, at least a part of the functionality described for the controllermay be implemented by the external device. The external device may be or may include any type of computing system or computing device, such as a desktop computer, notebook or laptop computer, pad or tablet device, smartphone, hand-held computing device, game controller, game system, and so on.
2060 2060 2060 2060 2060 2060 2060 2060 2060 In various embodiments, controllermay be a uniprocessor system including one processor, or a multiprocessor system including several processors (e.g., two, four, eight, or another suitable number). Controllermay include central processing units (CPUs) configured to implement any suitable instruction set architecture, and may be configured to execute instructions defined in that instruction set architecture. For example, in various embodiments controllermay include general-purpose or embedded processors implementing any of a variety of instruction set architectures (ISAs), such as the x86, PowerPC, SPARC, RISC, or MIPS ISAs, or any other suitable ISA. In multiprocessor systems, each of the processors may commonly, but not necessarily, implement the same ISA. Controllermay employ any microarchitecture, including scalar, superscalar, pipelined, superpipelined, out of order, in order, speculative, non-speculative, etc., or combinations thereof. Controllermay include circuitry to implement microcoding techniques. Controllermay include one or more processing cores each configured to execute instructions. Controllermay include one or more levels of caches, which may employ any size and any configuration (set associative, direct mapped, etc.). In some embodiments, controllermay include at least one graphics processing unit (GPU), which may include any suitable graphics processing circuitry. Generally, a GPU may be configured to render objects to be displayed into a frame buffer (e.g., one that includes pixel data for an entire frame). A GPU may include one or more graphics processors that may execute graphics software to perform a part or all of the graphics operation, or hardware acceleration of certain graphics operations. In some embodiments, controllermay include one or more other components for processing and rendering video and/or images, for example image signal processors (ISPs), coder/decoders (codecs), etc.
2070 Memorymay include any type of memory, such as dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate (DDR, DDR2,DDR3, etc.) SDRAM (including mobile versions of the SDRAMs such as mDDR3, etc., or low power versions of the SDRAMs such as LPDDR2, etc.), RAMBUS DRAM (RDRAM), static RAM (SRAM), etc. In some embodiments, one or more memory devices may be coupled onto a circuit board to form memory modules such as single inline memory modules (SIMMs), dual inline memory modules (DIMMs), etc. Alternatively, the devices may be mounted with an integrated circuit implementing system in a chip-on-chip configuration, a package-on-package configuration, or a multi-chip module configuration. In some embodiments DRAM may be used as temporary storage of images or video for processing, but other storage options may be used to store processed data, such as Flash or other “hard drive” technologies.
2000 2060 2050 2020 2000 2020 2060 2020 2000 2000 1 15 FIGS.through In some embodiments, devicemay include one or more sensors that collect information about the user's environment (video, depth information, lighting information, etc.). The sensors may provide the information to the controllerof the XR system. In some embodiments, the sensors may include, but are not limited to, at least one visible light camera (e.g., an RGB video camera), ambient light sensors, and at least on eye tracking camera. In some embodiments, devicemay also include one or more IR light sources; light from the light sources reflected off the eye may be captured by the eye tracking camera. Gaze tracking algorithms implemented by controllermay process images or video of the eye captured by the camerato determine eye pose and gaze direction. In addition, one or more of the methods as illustrated inmay be implemented in deviceto provide camera focusing in video passthrough mode for the device.
2000 2020 In some embodiments, devicemay be configured to render and display frames to provide an augmented or mixed reality (MR) view for the user based at least in part according to sensor inputs, including input from the eye tracking camera. The MR view may include renderings of the user's environment, including renderings of real objects in the user's environment, based on video captured by one or more video cameras that capture high-quality, high-resolution video of the user's environment for display. The MR view may also include virtual content (e.g., virtual objects, virtual tags for real objects, avatars of the user, etc.) generated by the XR system and composited with the displayed view of the user's real environment.
A real environment refers to an environment that a person can perceive (e.g., see, hear, feel) without use of a device. For example, an office environment may include furniture such as desks, chairs, and filing cabinets; structural items such as doors, windows, and walls; and objects such as electronic devices, books, and writing instruments. A person in a real environment can perceive the various aspects of the environment, and may be able to interact with objects in the environment.
An extended reality (XR) environment, on the other hand, is partially or entirely simulated using an electronic device. In an XR environment, for example, a user may see or hear computer generated content that partially or wholly replaces the user's perception of the real environment. Additionally, a user can interact with an XR environment. For example, the user's movements can be tracked and virtual objects in the XR environment can change in response to the user's movements. As a further example, a device presenting an XR environment to a user may determine that a user is moving their hand toward the virtual position of a virtual object, and may move the virtual object in response. Additionally, a user's head position and/or eye gaze can be tracked and virtual objects can move to stay in the user's line of sight.
Examples of XR include augmented reality (AR), virtual reality (VR) and mixed reality (MR). XR can be considered along a spectrum of realities, where VR, on one end, completely immerses the user, replacing the real environment with virtual content, and on the other end, the user experiences the real environment unaided by a device. In between are AR and MR, which mix virtual content with the real environment.
VR generally refers to a type of XR that completely immerses a user and replaces the user's real environment. For example, VR can be presented to a user using a head mounted device (HMD), which can include a near-eye display to present a virtual visual environment to the user and headphones to present a virtual audible environment. In a VR environment, the movement of the user can be tracked and cause the user's view of the environment to change. For example, a user wearing a HMD can walk in the real environment and the user will appear to be walking through the virtual environment they are experiencing. Additionally, the user may be represented by an avatar in the virtual environment, and the user's movements can be tracked by the HMD using various sensors to animate the user's avatar.
AR and MR refer to a type of XR that includes some mixture of the real environment and virtual content. For example, a user may hold a tablet that includes a camera that captures images of the user's real environment. The tablet may have a display that displays the images of the real environment mixed with images of virtual objects. AR or MR can also be presented to a user through an HMD. An HMD can have an opaque display, or can use a see-through display, which allows the user to see the real environment through the display, while displaying virtual content overlaid on the real environment.
a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and determine first focus distance information based on gaze information from the gaze tracking subsystem and a depth map; determine second focus distance information based on vergence of left and right gaze vectors as determined by the gaze tracking subsystem; determine a focus distance from the first focus distance information and the second focus distance information; and direct the camera to focus at the focus distance. a controller comprising one or more processors configured to: Clause 1. A device, comprising: model the first focus distance information and the second focus distance information as probability distance functions (PDFs); and determine the focus distance from the two PDFs. Clause 2. The device as recited in clause 1, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: Clause 3. The device as recited in clause 2, wherein the PDF corresponding to the first focus distance information indicates two or more possible distances, and wherein, to determine the focus distance from the two PDFs, the controller is configured to select one of the two or more distances that most closely corresponds to a distance indicated by the PDF corresponding to the second focus distance information. collect first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; apply a linear regression function to the collected first and second focus distance observations to generate calibrated vergence distances; and determine the focus distance from the calibrated vergence distances. Clause 4. The device as recited in clause 1, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: Clause 5. The device as recited in clause 4, wherein said collect first focus distance observations, said collect second focus distance observations, and said apply a linear regression function are performed during an enrollment process for the device. Clause 6. The device as recited in clause 4, wherein the observations are based on real objects in the scene that are imaged by the camera and displayed by the display subsystem or virtual objects rendered by the controller and displayed by the display subsystem. collect first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; train a model based on the collected first and second focus distance observations; and input at least one focus distance observation to the model, wherein the focus distance is output by the model in response to the input. Clause 7. The device as recited in clause 1, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: Clause 8. The device as recited in clause 7, wherein said collect first focus distance observations, said collect second focus distance observations, and said train a model are performed during an enrollment process for the device. display one or more targets at known ground truth positions; direct a user to fixate on at least one of the one or more targets; record left and right gaze vectors for the user while fixated on the targets; compute intersection points of the left and right gaze vectors and compare the intersection points with the ground truth positions of respective targets; store results of the comparison as calibrated vergence distances; and determine the focus distance from the calibrated vergence distances. Clause 9. The device as recited in clause 1, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: Clause 10. The device as recited in clause 1, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and determine a focus distance based on vergence of left and right gaze vectors as determined by the gaze tracking subsystem; direct the camera to focus at the focus distance; determine a metric from an image captured by the camera; adjust the focus distance if the metric is not within a specified range; and repeat said direct the camera, said determine a metric, and said adjust the focus distance until the metric is within the specified range. a controller comprising one or more processors configured to: Clause 11. A device, comprising: Clause 12. The device as recited in clause 11, wherein the metric is an image sharpness metric. Clause 13. The device as recited in clause 11, wherein the metric is disparity between left and right subpixels in one or more focus pixels. Clause 14. The device as recited in clause 11, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and estimate a gaze vector from gaze information captured by the gaze tracking subsystem; determine an intersection of the gaze vector with a depth map of the scene; determine a focus distance from an object or surface at the intersection of the gaze vector with the depth map; and direct the camera to focus at the determined focus distance. a controller comprising one or more processors configured to: Clause 15. A device, comprising: Clause 16. The device as recited in clause 15, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and estimate gaze vectors for left and right eyes from gaze information captured by the gaze tracking subsystem; determine vergence of the gaze vectors in the scene; determine a focus distance from the vergence of the gaze vectors; and direct the camera to focus at the determined focus distance. a controller comprising one or more processors configured to: Clause 17. A device, comprising: Clause 18. The device as recited in clause 17, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; at least one camera configured to capture images of a scene; and direct the display subsystem to display images including virtual content rendered from frames of the scene captured at the first focus distance; receive a signal that the camera is to be switched to a second focus distance; in response to the signal, direct the camera to focus at the second focus distance; and direct the display subsystem to display images including virtual content rendered from frames of the scene captured at the second focus distance. a controller comprising one or more processors configured to: direct the camera to focus at a first focus distance; Clause 19. A device, comprising: Clause 20. The device as recited in clause 19, wherein the first focus distance corresponds to a far focus mode of the camera, and wherein the second focus distance corresponds to a near focus mode of the camera. Clause 21. The device as recited in clause 20, wherein the far focus mode focuses the camera at 0.7 meters or more, and wherein the near focus mode focuses the camera at 0.4meters or less. Clause 22. The device as recited in clause 19, wherein the signal is generated in response to user input to the device manually changing a focus mode of the camera from the first focus distance to the second focus distance. Clause 23. The device as recited in clause 19, wherein the signal is generated in response to a user interaction with the device that indicates that a focus mode of the camera is to be switched from the first focus distance to the second focus distance. Clause 24. The device as recited in clause 19, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene to be displayed by the display subsystem; and a controller; a first focusing mode in which the controller directs the camera to focus at different distances determined from gaze information from the gaze tracking subsystem and a depth map combined with vergence of left and right gaze vectors determined by the gaze tracking subsystem; and a second focusing mode in which the controller directs the camera to focus at either a near focus distance or a far focus distance in response to input to the device indicating that the second focusing mode is to be used. wherein the device is configured to support two focusing modes for the camera: Clause 25. A device, comprising: Clause 26. The device as recited in clause 25, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene to be displayed by the display subsystem; and direct the camera to focus at a focus distance determined from gaze information from the gaze tracking subsystem and a depth map combined with vergence of left and right gaze vectors determined by the gaze tracking subsystem; determine an image sharpness metric from an image captured by the camera; adjust the focus distance if the image sharpness metric is not within a specified range; and repeat said direct the camera, said determine an image sharpness metric, and said adjust the focus distance until the image sharpness metric is within the specified range. a controller configured to: Clause 27. A device, comprising: Clause 28. The device as recited in clause 25, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content to an eye; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and determine a region of interest (ROI) based on an intersection of a gaze vector as determined by the gaze tracking subsystem with a focus pixel sensor image; direct the camera to focus at the ROI; determine sub-pixel disparity for the ROI in the focus pixel sensor image; adjust the focus distance if disparity is detected between left and right subpixels in the region of interest; and repeat said direct the camera, said determine a metric, and said adjust the focus distance until the sub-pixel disparity is within a specified range. a controller comprising one or more processors configured to: Clause 29. A device, comprising: Clause 30. The device as recited in clause 29, wherein the device is a head-mounted device (HMD) of an extended reality (XR) system. determining first focus distance information based on gaze information from a gaze tracking subsystem and a depth map; determining second focus distance information based on vergence of left and right gaze vectors as determined by the gaze tracking subsystem; determining a focus distance from the first focus distance information and the second focus distance information; and directing a camera to focus at the focus distance. performing, by a controller comprising one or more processors: Clause 31. A method, comprising: modeling the first focus distance information and the second focus distance information as probability distance functions (PDFs); and determining the focus distance from the two PDFs. Clause 32. The method as recited in clause 31, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises: collecting first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collecting second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; applying a linear regression function to the collected first and second focus distance observations to generate calibrated vergence distances; and determining the focus distance from the calibrated vergence distances. Clause 33. The method as recited in clause 31, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises: collecting first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collecting second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; training a model based on the collected first and second focus distance observations; and inputting at least one focus distance observation to the model, wherein the focus distance is output by the model in response to the input. Clause 34. The method as recited in clause 31, wherein determining a focus distance from the first focus distance information and the second focus distance information comprises: Clause 35. The method as recited in clause 31, wherein the controller, display subsystem, gaze tracking subsystem, and camera are components of a head-mounted device (HMD) of an extended reality (XR) system. determining a focus distance based on vergence of left and right gaze vectors as determined by a gaze tracking subsystem; directing a camera to focus at the focus distance; determining a sharpness from an image captured by the camera; adjusting the focus distance if the sharpness metric is not within a specified range; and repeating said direct the camera, said determine a metric, and said adjust the focus distance until the sharpness metric is within the specified range. performing, by a controller comprising one or more processors: Clause 36. A method, comprising: Clause 37. The method as recited in clause 36, wherein the controller, display subsystem, gaze tracking subsystem, and camera are components of a head-mounted device (HMD) of an extended reality (XR) system. directing a camera to focus at a first focus distance; directing a display subsystem to display images including virtual content rendered from frames of the scene captured at the first focus distance; receiving a signal that the camera is to be switched to a second focus distance; in response to the signal, directing the camera to focus at the second focus distance; and directing the display subsystem to display images including virtual content rendered from frames of the scene captured at the second focus distance; performing, by a controller comprising one or more processors: wherein the first focus distance corresponds to a far focus mode of the camera, and wherein the second focus distance corresponds to a near focus mode of the camera. Clause 38. A method, comprising: Clause 39. The method as recited in clause 38, wherein the signal is generated in response to user input manually changing a focus mode of the camera from the first focus distance to the second focus distance. Clause 40. The method as recited in clause 38, wherein the controller, display subsystem, gaze tracking subsystem, and camera are components of a head-mounted device (HMD) of an extended reality (XR) system. a display subsystem configured to display virtual content; a gaze tracking subsystem; at least one camera configured to capture images of a scene; and determine first focus distance information based on gaze information from the gaze tracking subsystem and a depth map; determine second focus distance information based on vergence of left and right gaze vectors as determined by the gaze tracking subsystem; determine a focus distance from the first focus distance information and the second focus distance information; and direct the camera to focus at the focus distance. a controller comprising one or more processors configured to: a head-mounted device (HMD), comprising Clause 41. A system, comprising: model the first focus distance information and the second focus distance information as probability distance functions (PDFs); and determine the focus distance from the two PDFs. Clause 42. The system as recited in clause 41, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: collect first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; apply a linear regression function to the collected first and second focus distance observations to generate calibrated vergence distances; and determine the focus distance from the calibrated vergence distances. Clause 43. The system as recited in clause 41, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: collect first focus distance observations based on the gaze information from the gaze tracking subsystem and the depth map; collect second focus distance observations based on the vergence of left and right gaze vectors as determined by the gaze tracking subsystem; train a model based on the collected first and second focus distance observations; and input at least one focus distance observation to the model, wherein the focus distance is output by the model in response to the input. Clause 44. The system as recited in clause 41, wherein, to determine a focus distance from the first focus distance information and the second focus distance information, the controller is configured to: Clause 45. The system as recited in clause 41, wherein the system is an extended reality (XR) system. The following clauses describe various examples of embodiments consistent with the description provided herein.
The methods described herein may be implemented in software, hardware, or a combination thereof, in different embodiments. In addition, the order of the blocks of the methods may be changed, and various elements may be added, reordered, combined, omitted, modified, etc. Various modifications and changes may be made as would be obvious to a person skilled in the art having the benefit of this disclosure. The various embodiments described herein are meant to be illustrative and not limiting. Many variations, modifications, additions, and improvements are possible. Accordingly, plural instances may be provided for components described herein as a single instance. Boundaries between various components, operations and data stores are somewhat arbitrary, and particular operations are illustrated in the context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within the scope of claims that follow. Finally, structures and functionality presented as discrete components in the example configurations may be implemented as a combined structure or component. These and other variations, modifications, additions, and improvements may fall within the scope of embodiments as defined in the claims that follow.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
September 17, 2024
August 11, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.