Patentable/Patents/US-20260270561-A1
US-20260270561-A1

Method and Device for Capturing Images of a Moving Subject

PublishedSeptember 10, 2026
Assigneenot available in USPTO data we have
Technical Abstract

capture capture A method is provided for capturing images of a region of interest of a moving subject. The method includes steps of computation, repeated at a computation frequency, of the three-dimensional position of the region of interest in the capture volume; estimation of a three-dimensional path of the region of interest from the computed positions, the path being defined by a parametric path model combining a periodic and linear temporal component; determination of a position of the region of interest at least at one future time tfrom the estimated path; and capture of an image of the region of interest with the capture camera pointing in the direction of the determined position of the region of interest at the time t.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

103 computation, repeated at a computation frequency, of a three-dimensional position of the region of interest, delivering as output a sequence of three-dimensional positions of the region of interest, on which an instantaneous observable position signaldepends, from a series of acquired images in which the region of interest is detected and located; extraction, at an application frequency, of a periodic observable signalfrom the instantaneous observable position signal; eyes periodic linear estimation of a three-dimensional path (P(t)) of the region of interest in the capture volume from the computed positions, the path being defined, in at least one dimension, by a parametric path model combining a periodic temporal component (P(t)) and a linear temporal component (P(t)); capture determination of a position of the region of interest at least at one future time tfrom the estimated three-dimensional path of the region of interest; capture capture capture, at the time t, of an image of the region of interest by a capture camera with the capture camera pointing in the direction of the determined position of the region of interest at the time t. . A method for capturing images of a region of interest of a moving subject () within a capture volume, the method comprising the following steps:

2

claim 1 . The method as claimed in, wherein the subject is a person, the region of interest being a portion of a face of the moving subject, in particular an eye, preferably an iris, or being a visual representation worn by the subject, such as a two-dimensional bar code.

3

claim 1 eyes . The method as claimed in, wherein the step of estimation of the path (P(t)) comprises implementation, for the detected region of interest of the subject, of an extended Kalman filter of the parametric path model in said at least one dimension, this comprising a phase of initialization of the extended Kalman filter, a phase of prediction by the extended Kalman filter and a phase of update of the extended Kalman filter at a sampling frequency.

4

claim 1 . The method as claimed in, comprising a step of application, at the application frequency, of a band-pass filter in said at least one dimension to the instantaneous observable position signalwith a view to extracting the periodic observable signal.

5

claim 3 . The method as claimed in, wherein the extended Kalman filter is based on a measurement vector containing at least two observers, having as first observer the instantaneous observable position signal, which is made up of the sequence of computed positions of the region of interest, and as second observer the periodic observable signal.

6

claim 3 . The method as claimed in, wherein a state vector of the extended Kalman filter average path contains five states, namely an average position of the region of interest in said at least one dimension (P(t)), a rate of change in the average position of the region of interest in said at least one dimension, an amplitude (Amplitude) of oscillation of the region of interest around the average position, a phase (Φ(t)) of oscillation around the average position and an angular frequency (Φ(t)) of oscillation of the average position of the region of interest.

7

claim 3 0 an amplitude parameter (Amplitude); 0 a phase parameter (Φ); an angular-frequency parameter (ω); a rate-of-change parameter (drift); 0 an intercept parameter (P); said initialization parameters being determined from all or part of the sequence of computed three-dimensional positions of the region of interest. . The method as claimed in, wherein the phase of initialization of the filter comprises determining an initial state vector (X(t)) through determination of initialization parameters comprising:

8

claim 7 . The method as claimed in, wherein said part of the sequence of positions comprises the positions computed for said region of interest on the basis of successively acquired images, the first image of which is the one in which the region of interest is detected a first time, and the following images of which are those in which the region of interest is detected, until at least two local extrema are detected.

9

claim 3 s drift . The method as claimed in, wherein a covariance matrix of the process noise (Q) of the extended Kalman filter depends on the sampling period (T) of the extended Kalman filter and on at least one variance (q) of the rate of change angular frequency α in the average position of the region of interest in said at least one dimension, a variance (q) of the angular frequency of oscillation (Φ(t)) of the region of interest around the average position of the region of interest or a variance (q) of the amplitude of oscillation (Amplitude) of the region of interest around the average position of the region of interest.

10

claim 3 tracking tracking . The method as claimed in, wherein the estimation of the path of the region of interest at an arbitrary future time tis obtained by computing a tangent to the modeled path at said arbitrary future time tand the parameters of which result from the update of the extended Kalman filter on the basis of images acquired prior to the estimation, in particular with the modeled path written in terms of position pos and speed spd as: measure position with tthe time at which the last context image used to estimate the path is acquired, the tangent to the parametric model being written:

11

claim 1 . The method as claimed in, wherein the capture camera is pointed in the direction of the position of the region of interest determined depending on the three-dimensional path of the region of interest at a frequency higher than the computation frequency.

12

claim 1 . A computer program comprising instructions configured to implement each of the steps of the method as claimed inwhen said program is executed on a computer.

13

claim 1 . A non-transitory computer-readable medium that stores code instructions of a computer program for causing a computer or microprocessor to execute the method as claimed in.

14

a context camera system; a capture camera; computation, repeated at a computation frequency, of a three-dimensional position of the region of interest, delivering as output a sequence of three-dimensional positions of the region of interest, on which an instantaneous observable position signaldepends, from a series of images acquired by the context camera system and in which the region of interest is detected and located; extraction, at an application frequency, of a periodic observable signalfrom the instantaneous observable position signal; eyes periodic linear estimation of a three-dimensional path (P(t)) of the region of interest in the capture volume from the computed positions, the path being defined, in at least one dimension, by a parametric path model combining a periodic temporal component (P(t)) and a linear temporal component (P(t)); capture determination of a position of the region of interest at least at one future time tfrom the estimated three-dimensional path of the region of interest; capture capture capture, at the time t, of an image of the region of interest with the capture camera pointing in the direction of the determined position of the region of interest at the time t. a processor configured for the following steps: . A device for capturing images of a region of interest of a moving subject within a capture volume, the device comprising:

15

claim 14 . The device as claimed in, wherein the capture camera is rotatable so as to point in the direction of the determined position of the region of interest in said at least one dimension.

16

claim 14 . The device as claimed in, wherein the subject is a person, the region of interest being a portion of a face of the moving subject, in particular an eye, preferably an iris, or being a visual representation worn by the subject, such as a two-dimensional bar code.

17

claim 14 eyes . The device as claimed in, wherein the estimation of the path (P(t)) comprises implementation, for the detected region of interest of the subject, of an extended Kalman filter of the parametric path model in said at least one dimension, this comprising a phase of initialization of the extended Kalman filter, a phase of prediction by the extended Kalman filter and a phase of update of the extended Kalman filter at a sampling frequency.

18

claim 17 . The device as claimed in, wherein the extended Kalman filter is based on a measurement vector containing at least two observers, having as first observer the instantaneous observable position signal, which is made up of the sequence of computed positions of the region of interest, and as second observer the periodic observable signal.

19

claim 17 . The device as claimed in, wherein a state vector of the extended Kalman filter average path contains five states, namely an average position of the region of interest in said at least one dimension (P(t)), a rate of change in the average position of the region of interest in said at least one dimension, an amplitude (Amplitude) of oscillation of the region of interest around the average position, a phase (Φ(t)) of oscillation around the average position and an angular frequency (Φ(t)) of oscillation of the average position of the region of interest.

20

claim 14 . The device as claimed in, the processor being further configured to perform a step of application, at the application frequency, of a band-pass filter in said at least one dimension to the instantaneous observable position signalwith a view to extracting the periodic observable signal.

Detailed Description

Complete technical specification and implementation details from the patent document.

The present invention relates to the field of capturing images of a region of interest located on a moving subject. More particularly, the invention relates to tracking and obtaining a sharp image of the region of interest located on a moving person.

In a first example, it may be a question of recognizing the irises of a person passing in front of the image-capturing system, in the case of biometric recognition of a moving person. In a second example, it may also be a question of reading information, such as a two-dimensional code or QR code, of a badge worn by a moving person. Other examples of technical contexts may benefit from the invention.

Systems allowing the region of interest to be captured when the subject is stationary already exist. The capture volume is the region of space considered when capturing the image of the region of interest. It is the region in which the subject is able to move while images of the region are able to be captured. If the subject exits the capture volume, they are no longer visible to the image-capturing system.

When the subject is stationary, the image-capturing system may operate according to the following principle. In a first step, the subject is detected and located within the capture volume. Next, the image-capturing system waits for the subject to stabilize. When the subject is stable, an exact step of locating the region of interest makes it possible to obtain the coordinates thereof in three dimensions. These coordinates then allow the capture camera, which is equipped with motors allowing its orientation, to be positioned in the exact direction of the targeted region of interest. A succession of images is then captured until the obtained quality is sufficient for the application, typically iris recognition or read-out of information. It will be noted that the envisioned applications require an image of high resolution to be captured, irrespectively of whether it is for iris recognition or read-out of information.

When the subject is moving, capturing sharp images of the region of interest becomes more difficult. Specifically, the movement of the subject may generate blurry images. In addition, due to the required resolution and sensitivity, the captured images have a very shallow depth of field. An inaccuracy in the distance of the region of interest then also results in capture of a blurry image. Furthermore, it has also been observed that when capturing an image of an iris at short range, and in particular at less than one meter, pointing becomes tricky, since the movement of the person can no longer be compensated for by the field of the camera, and the target exits the field of the capture camera. Similarly, at short range, it has been observed that the temporal latency between detection of the eyes and the control of movement of the capture camera by its motors can have a substantial impact on the accuracy of the estimated position. These drawbacks may result in failure when attempting to track at short range or high walking speed.

The invention aims to solve all or some of these problems by proposing a method for capturing images of a region of interest of a moving subject that allows acquisition, in a reliable manner, of a sharp image of the targeted region of interest of the moving subject, in particular at short range (for example at less than one meter, and in particular up to 70 cm from the capture camera) including when the subject is walking at high speed (for example at more than 1 m/s, and in particular up to 1.5 m/s).

computation, repeated at a computation frequency, of a three-dimensional position of the region of interest, delivering as output a sequence of three-dimensional positions of the region of interest, on which an instantaneous observable position signal depends, from a series of acquired images in which the region of interest is detected and located; extraction, at an application frequency, of a periodic observable signal from the instantaneous observable position signal; estimation of a three-dimensional path of the region of interest in the capture volume from the computed positions, the path being defined, in at least one dimension, by a parametric path model combining a periodic temporal component and a linear temporal component; capture determination of a position of the region of interest at least at one future time tfrom the estimated three-dimensional path of the region of interest; capture capture capture, at the time t, of an image of the region of interest by a capture camera with the capture camera pointing in the direction of the determined position of the region of interest at the time t. According to one aspect of the invention, a method for capturing images of a region of interest of a moving subject within a capture volume is proposed, the method comprising the following steps:

This method allows a sharp image of the region of interest of a moving subject to be captured, including at short range, with a short tracking time during which a sequence of three-dimensional positions of the region of interest is obtained after detection and location. Furthermore, this method makes it possible to compensate for latencies. Thus, the number of images of the subject that are captured by the capture camera is reduced, capture by the capture camera of a single image of the subject possibly even sufficing, due to the obtained image quality, for the application, in particular biometric application, the capture of a succession of images of the subject by the capture camera no longer necessarily being required. Furthermore, this method allows multiple subjects to be tracked in parallel.

In particular, the motion of the subject is a walking motion, the speed of which is in particular less than 2 m/s, and particularly between 1 and 1.5 m/s, the method therefore being suitable for average walking speeds but also hurried walking speeds.

Advantageously, the acquired images are images of a detection volume, which are acquired by a context camera system, and which are also called context images, this allowing tracking by the context camera, during which tracking a sequence of three-dimensional positions of the region of interest in the detection volume is obtained after detection and location. Similarly, if multiple subjects are present in the detection volume and appear in the context images, one region of interest per subject will be tracked.

Advantageously, if the acquisitions are carried out at a given acquisition frequency, the computation frequency is in particular less than or equal to said acquisition frequency.

Advantageously, the estimation is carried out at an estimation frequency, preferably equal to the computation frequency.

Advantageously, images are captured by a capture camera in the capture volume at a capture frequency.

Advantageously, the detection volume, which corresponds to the segment of the field of the context camera system in which the acquired image quality is sufficient to carry out object detection (for example of depth greater than 2 meters from the context camera system, and preferably 2.50 meters), and the capture volume, which corresponds to the segment of the field of the capture camera in which the captured image quality is sufficient to achieve recognition of the region of interest, have a common portion, this making it possible to work with a detection volume having a depth greater than the depth of the capture volume, which is in particular constrained by the high resolution required for iris acquisition in the context of biometric recognition (for example depth between 1.50 meters and 0.50 meters from the capture camera in the case of iris recognition). The large depth of the detection volume allows tracking such that the model will have already converged before the subject enters the capture volume, given that initialization for example occurs on a half step movement, i.e. of about 50 centimeters.

Preferably, the instantaneous observable position signal is made up of the sequence of positions computed for the region of interest.

Advantageously, the periodic temporal component of the path model comprises an amplitude parameter, a phase parameter and an angular-frequency parameter, allowing a periodic sinusoidal temporal component to be represented. In a manner equivalent to this spectral angular-frequency decomposition, a frequency decomposition is possible.

Advantageously, the linear temporal component of the path model comprises a rate-of-change parameter and an intercept parameter.

Advantageously, one of said at least one dimension is the vertical dimension (y), this allowing the height of the region of interest of a moving subject to be accurately modeled, in particular by taking into account its up-down rhythm, this being especially important to pointing the capture camera in the direction of the determined position of the region of interest, especially when the subject is walking.

Advantageously, one of said at least one dimension is the lateral dimension (x), this allowing the horizontal left-right lateral swaying motion (motion in the x-direction) of the region of interest of a moving subject to be accurately modeled, especially when the subject is walking.

Similarly, said at least one dimension may designate both the aforementioned dimensions because the step, i.e. the gait, of a subject includes not only an up-down vertical motion (in the y-direction) but also a left-right swaying motion (in the x-direction).

Advantageously, one of said at least one dimension is the longitudinal dimension (z), this allowing the horizontal longitudinal depthwise motion (motion in the z-direction) of the region of interest of a moving subject to be accurately modeled, especially when the subject is walking. Specifically, to a lesser extent the path in the depth dimension (z) may also be thus modeled, this achieving a better estimation than a linear approximation alone because the gait of a subject follows a specific rhythm that, for example, varies depending on the supporting leg, the motion resulting from which, in particular when the person limps, is not necessarily symmetrical, creating a back-and-forth rocking effect with each step.

Thus, said at least one dimension may advantageously designate all three of the aforementioned dimensions so as to allow the location of the region of interest in said three dimensions to be accurately estimated.

capture Advantageously, pointing of the capture camera is accompanied by adjustment of the focus of the capture camera as a function of the distance between the capture camera and the determined position of the region of interest at the time t, this making it possible to improve the sharpness of the obtained image, in particular if the depth of field of the lens of the capture camera is not sufficiently wide. Advantageously, the adjustment and pointing may be carried out simultaneously.

capture Advantageously, the focus of the capture camera is adjusted depending on the distance between the capture camera and the determined position of the region of interest at the time tplus a focus factor that varies in set or variable steps in an interval, in particular one centered on zero, this allowing a focus ramp to be achieved.

Advantageously, the focus factor varies linearly between the extreme values of the zero-centered interval.

Advantageously, the focus factor varies sinusoidally between the extreme values of the zero-centered interval.

Advantageously, the focus factor varies in increments corresponding to the depth of field between the extreme values of the zero-centered interval.

In one embodiment, the subject is a person, the region of interest being a portion of a face of the moving subject, including an eye, preferably an iris, or being a visual representation (pictogram) worn by the subject, such as a two-dimensional bar code.

Advantageously, the capture frequency of the capture camera is greater than the frequency of acquisition by the context camera system.

In one embodiment, the step of estimation of the path comprises implementation, for the detected region of interest of the subject, of an extended Kalman filter of the parametric path model in said at least one dimension, this comprising a phase of initialization of the extended Kalman filter, a phase of prediction by the extended Kalman filter and a phase of update of the extended Kalman filter at a sampling frequency. The extended Kalman filter makes it possible to linearize locally the non-linear component of the parametric path model for the region of interest in question. The extended Kalman filter makes it possible to simultaneously make a prediction (future value) and denoise the models with non-linear equations, something that is particularly useful when it is considered that the instantaneous observable position signal may be decomposed into a combination of a linear function, of a sinusoidal function and of measurement noise. In particular, if a plurality of subjects are detected, the step of estimation of the path comprises implementation, for the detected region of interest of each subject, of an extended Kalman filter of the parametric path model in said at least one dimension.

The computed or determined positions and said estimated path are preferably in, or transposed to, a frame of reference belonging to the capture camera.

capture In particular, the determination of the position of the region of interest at least at one future time tfrom the path estimated through implementation of an extended Kalman filter is carried out only if said Kalman filter has converged.

Advantageously, the path estimation frequency is the same as the sampling frequency of the extended Kalman filter, which frequency is also called the update frequency of the extended Kalman filter, and less than or equal to a frequency of acquisition of context images by the context camera system, which itself is preferably greater than or equal to said computation frequency, i.e. the frequency of computation of the three-dimensional position of the region of interest from the series of acquired context images.

Advantageously, the three-dimensional path of the region of interest is estimated at a frequency equal to the frequency of computation of the three-dimensional position of the region of interest, this allowing the computations to be synchronized. As a variant, the sampling frequency of the extended Kalman filter may be variable, for example in the case of a context camera system comprising a multiplicity of sensors the data of which are fused, and that have different frequencies, the position computations hence not necessarily occurring at synchronized times, and the filter is preferably updated with each new computed position, such a device and method in particular making it possible to limit concealment in the detection volume, by virtue of the fusion of sensor data.

Advantageously, the frequency of computation of the position varies over time, this making it possible to adapt to the availability of sensor measurements, in particular if a target is momentarily concealed—for example, in the context of iris acquisition, the person could turn their head, which would make their eyes disappear from the field of acquisition, or a change in the exposure time of the context camera system could occur depending on lighting conditions. Similarly, this also makes it possible to mitigate image loss due to overload of the central processing unit of the data-processing device implementing the invention. For example, in case of non-detection of the region of interest in an image, the frequency of computation of the position may be decreased. Thus, the update frequency of the extended Kalman filter is variable, depending on whether or not the region of interest is detected in the acquired images, the extended Kalman filter not being updated in computation cycles in which the region of interest has not been detected in the acquired image in said cycle, thus avoiding computations.

In one embodiment, the method comprises a step of application, at an application frequency, of a band-pass filter in said at least one dimension to the instantaneous observable position signal with a view to extracting the periodic observable signal.

In particular, the periodic observable signal is centered on zero, this resulting from the application of the band-pass filter and achieving compatibility with the periodic temporal component of the zero-centered path model.

Advantageously, the update frequency of the extended Kalman filter is equal to the application frequency of the band-pass filter.

In particular, the application frequency of the band-pass filter is variable, this making it possible to cope with cases of concealment of the subject, i.e. when the three-dimensional position cannot be computed in certain acquired images. Advantageously, the band-pass filter is biquadratic, allowing customized and robust implementation.

Advantageously, the coefficients of the band-pass filter are determined experimentally, based on tests carried out on a sample of subjects (in particular at varied walking speeds) whose average oscillation frequency (in the y-direction) preferably varies between 1 and 2 Hz, by spectral analysis of the data thus obtained.

In one embodiment, the extended Kalman filter is based on a measurement vector containing at least two observers, having as first observer the instantaneous observable position signal, which is made up of the sequence of computed positions of the region of interest, and as second observer the periodic observable signal.

In one embodiment, a state vector of the extended Kalman filter contains five states, namely an average position of the region of interest in said at least one dimension, a rate of change in the average position of the region of interest in said at least one dimension, an amplitude of oscillation of the region of interest around the average position, a phase of oscillation around the average position and an angular frequency of oscillation of the average position of the region of interest. For example, in the case of application to an iris, in the case of a parametric path model in the dimension y only, the states listed above for example correspond to the height of the person's eyes, to the rate of change in this height (excluding any oscillation (“linear” component)), to the amplitude of the oscillation of the height of the eyes, to the phase of the sinusoid modeling this oscillation around the average height of the eyes and to the rate of change in this phase, i.e. the angular frequency of oscillation, in particular resulting from the pace of the steps of the person. This composition of the states makes it possible to improve accuracy by increasing fidelity to parameters describing gait. Equivalently, the angular frequency and phase may be replaced by a frequency and phase.

an amplitude parameter; a phase parameter; an angular-frequency parameter; a rate-of-change parameter; an intercept parameter;said initialization parameters being determined from all or part of the sequence of computed three-dimensional positions of the region of interest. In particular, for said region of interest, a single initialization phase is sufficient, including in the case of temporary concealment of the subject in one or more images. In one embodiment, the phase of initialization of the extended Kalman filter comprises determining an initial state vector through determination of initialization parameters comprising:

Advantageously, said part of the sequence comprises at least two positions computed on the basis of images acquired successively by the context camera system, and for example comprises the first n computed three-dimensional positions of the region of interest, with n in particular being greater than or equal to 4, and for example 8, particularly for a subject whose step frequency is substantially one hertz.

In one embodiment, said part of the sequence of positions comprises the positions computed for said region of interest on the basis of successively acquired images, the first image of which is the one in which the region of interest is detected a first time, and the following images of which are those in which the region of interest is detected, until at least two local extrema are detected, this allowing the end of a step movement to be reached and initialization to occur on at least one half step movement, the middle of the step movement for example being considered to correspond to a local maximum and the end of the step movement to correspond to a local minimum (or vice versa). This embodiment therefore allows a robust initialization appropriate for a multidimensional path model.

In one embodiment, a covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and on at least one variance of the rate of change in the average position of the region of interest in said at least one dimension, a variance of the angular frequency of oscillation of the region of interest around the average position of the region of interest or a variance of the amplitude of oscillation of the region of interest around the average position of the region of interest. Thus, the covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and on at least one variance of the representative rate of the moving subject, of the period of oscillation representative of the motion of the moving subject or of the amplitude of oscillation representative of the motion of the moving subject. Said variances are preferably pre-calibrated empirically (for example on the basis of a statistical pre-study) and/or depend on positions of the region of interest that are computed before the prediction phase.

Preferably, the covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and on all three variances.

tracking tracking In one embodiment, the estimation of the path of the region of interest at an arbitrary future time tis obtained by computing a tangent to the modeled path at said arbitrary future time tand the parameters of which result from the update of the extended Kalman filter on the basis of images acquired prior to the estimation, in particular with the modeled path written in terms of position pos and speed spd as:

measure position with tthe time at which the last context image used to estimate the path is acquired, the tangent to the parametric model being written:

capture capture this making it possible, by applying t=t, to obtain the determination of the position of the region of interest at the future time t.

tracking The arbitrary future time tis advantageously an estimated time of receipt of the path by the real-time coprocessor, this making it possible to take into account the latency between the real time and the times of acquisition of the context images.

tracking In one embodiment, the capture camera is pointed in the direction of the position of the region of interest determined depending on the three-dimensional path of the region of interest at a frequency higher than the computation frequency, this allowing the pointing direction to be changed at a rate greater than the position computation frequency and real-time tracking to be carried out on the basis of the latest estimation of the path of the region of interest at t.

Advantageously, the context camera system comprises two cameras producing stereoscopic images.

Advantageously, the context camera system comprises a time-of-flight camera.

According to another aspect of the invention, a computer program comprising instructions configured to implement each of the steps of the method according to the invention when said program is executed on a computer is proposed.

According to another aspect of the invention, a removable or irremovable information storage means that is partially or totally readable by a computer or a microprocessor is proposed, said means containing code instructions of a computer program for executing each of the steps of the method according to the invention.

a context camera system; a capture camera; computation, repeated at a computation frequency, of a three-dimensional position of the region of interest, delivering as output a sequence of three-dimensional positions of the region of interest, on which an instantaneous observable position signal depends, from a series of images acquired by the context camera system and in which the region of interest is detected and located; extraction, at an application frequency, of a periodic observable signal from the instantaneous observable position signal; estimation of a three-dimensional path of the region of interest in the capture volume from the computed positions, the path being defined, in at least one dimension, by a parametric path model combining a periodic temporal component and a linear temporal component; capture determination of a position of the region of interest at least at one future time tfrom the estimated three-dimensional path of the region of interest; capture capture capture capture capture, at the time t, of an image of the region of interest with the capture camera pointing in the direction of the determined position of the region of interest at the time t; this allowing the capture camera to be pointed in the direction of the determined position of the region of interest at the time t, and an image to be captured by the capture camera at the time t, and having the same advantages as the method according to the invention. a processor configured for the following steps: According to another aspect of the invention, a device for capturing images of a region of interest of a moving subject within a capture volume is proposed, the device comprising:

In one embodiment, the capture camera is rotatable so as to point in the direction of the determined position of the region of interest in said at least one dimension.

Advantageously, the context camera system comprises two cameras producing stereoscopic images.

Advantageously, the context camera system comprises a time-of-flight camera.

Identical references have been used in all the figures to designate elements that are identical or similar, in their form or function.

103 100 103 100 The invention is applicable to various contexts. It may be a question of recognition of the irises of a moving personpassing in front of an image-capturing device, or even of recognition of a badge worn by a moving personpassing in front of an image-capturing device, for example with the aim of permitting or preventing said person from gaining access and/or of timestamping their entry. The subject may, for example, be passing through a security gate (such as in an airport, a virology laboratory or a nuclear power plant) or simply be moving about in a free space (such as a public space) and be tracked by the image-capturing device. Another example of embodiment of the invention relates to capture of a two-dimensional code or QR code (QR standing for Quick Response) placed on an object that is being moved about by a moving person, in particular one who is walking, in particular during delivery or handling. It is then a question of capturing images of the code of the object while it is being handled, typically in a warehouse. In each and every case, it is a question of capturing a sharp image of a region of interest located on or worn by a moving subject.

1 FIG. 103 100 103 100 The embodiment ofpertains to the context of recognition of the irises of a moving personpassing in front of an image-capturing deviceaccording to one embodiment of the invention, arranged in an airport. The acquisition in motion of the irises of the personallows the person to be recognized and authenticated upstream of the device, this for example allowing a security gate that would otherwise prevent access to be opened if the person is recognized as being permitted entry, and where appropriate a record of their entry to be kept, without the person having to stop in particular, making entry fluid and fast, the person being able to maintain their walking speed.

2 FIG. 100 101 102 102 103 104 102 101 104 103 illustrates the architecture of the deviceaccording to one embodiment of the invention. This figure illustrates the capture volumeequipped with a capture camera. The capture cameraadvantageously makes it possible to capture an image of a subjecthaving a region of interest. To do this, the capture camerais typically motorized to allow it to be oriented in space. In the example of embodiment, it is equipped with two motors one of which allows rotation in the horizontal plane and the other of which allows rotation in the vertical plane. The movement of these two motors makes it possible to point the capture camera in any direction within the capture volume. A third motor adjusts the focus of the capture camera, i.e. adapts the image capture to the distance of the subject from the capture camera, and more particularly to the distance of the region of interestof the subjectfrom the capture camera, this distance here corresponding to a depth.

100 106 106 102 The image-capturing deviceis controlled by a data-processing devicethat allows the movements of these motors to be controlled. Typically, but not necessarily, it is also this data-processing devicethat receives and processes the images received from the capture camera. The data-processing device is typically a computer, a tablet, a smart phone or any other device allowing a computer program responsible for controlling the camera, acquiring images and for the various steps of the method according to the invention to be executed.

102 105 105 103 104 106 106 103 104 In a reference embodiment, the capture camerais supplemented by a context camera system, in particular comprising one or more context cameras, the distance from the subject possibly being computed by triangulation using at least two context cameras. This context camera system makes it possible to capture images of the detection volume, which covers all or part of the capture volume, then to recognize a subjectin the acquired image before locating them, then to locate the region of interestof the subject. In certain embodiments, the context camera system is also controlled by the data-processing device. Alternatively, a dedicated device may be responsible for controlling the context camera system. In this case, this device communicates with the devicewith a view to capturing images. The context camera system may, as a variant, or in addition to a context camera, comprise a time-of-flight (ToF) sensor, a radar or any other sensor able to contribute to locating the subject. The context cameras and their optional ancillary sensors make it possible to compute the three-dimensional position of the subjectand more particularly of the region of interestof the subject in space, and in particular in the detection and capture volumes.

105 102 This estimation of the position of the region of interest is used to position and adjust the capture of images by the capture camera. However, the acquisition of images by the context cameras, their analysis to determine the three-dimensional position of the region of interest, the computation of the motor controls allowing the capture camerato be directed toward the computed position and the acquisition of images by the capture camera takes a certain time. There are therefore latencies between the computation of the position and the capture of the images of the region of interest. Due to the continuous movement of the subject, to these latencies and to the inherent inaccuracy of the positioning computation, obtaining a sharp capture image of the region of interest is a challenge.

103 104 According to the invention, provision is made to control the capture camera depending, not on the position of the region of interest computed by the context cameras, but on an estimation of where this position will be when an image of this region of interest is actually captured by the capture camera. The estimation of position is therefore an estimation of a future position of the region of interest at the time of its estimation. It is made based on computed actual positions allowing the path of the subjectto be estimated and therefore their position in the future, at the moment when the capture camera will capture the image, to be determined. By path what is meant is the position, as a function of time, that models the motion as a function of time—in other words, the path of a point is the geometrical curve that it describes as it moves through the frame of reference and it therefore provides information on position and speed over time in the frame of reference. To determine this future position, a three-dimensional path of the region of interestis estimated from the computed positions; the path is, in at least one dimension, defined by a parametric path model combining a periodic temporal component and a linear temporal component. It will thus be understood that the future position is determined with a high accuracy apposite to the movement of the subject, whether the latter is walking or in a wheelchair for example, by virtue of the combination of the two components. Specifically, if the person is in a wheelchair, the periodic component of the parametric model will be estimated to be zero.

Where appropriate, focusing may be carried out, the determined three-dimensional position then being modified by a focus factor in the dimension representing the distance of the region of interest from the capture camera. This focus factor represents the addition of a value that varies linearly between a negative minimum and a positive maximum around the value of zero and changes between each image capture. Adjusting the focus of the image captures slightly in this way increases the chances of obtaining at least one sharp image in a set of captured images.

Advantageously, the motors with which the capture camera is equipped are brushless motors, and in particular direct drive motors (allowing continuous motor movement). Brushless motors allow smooth, continuous tracking by the capture camera of the motion of the subject, without compromising image quality. When the motors are direct-drive motors, the absence of reduction gears specific to these motors eliminates the associated mechanical play. This would not be possible with stepper motors controlled using a stepped control method, as used in known systems.

3 FIG. 100 100 201 106 201 illustrates the hardware architecture of an image-capturing deviceaccording to one example of embodiment of the invention. The image-capturing deviceis mainly controlled by the main processor, which is preferably, but not necessarily, hosted in the data-processing device. This processoris typically operated by a conventional operating system such as Linux®, Windows® or MacOS®. These operating systems are not strictly real-time, potentially affecting the accuracy of the estimations, and this must be taken into account.

100 202 106 202 202 211 105 212 102 213 205 214 206 215 207 209 105 210 102 201 The image-capturing devicealso comprises a coprocessor, which is preferably, but not necessarily, hosted in the data-processing device. The main characteristic of this coprocessoris that it is real-time, i.e. able to execute certain tasks at precise predetermined times, in a guaranteed manner. The real-time coprocessoris responsible for triggeringthe context cameras, triggeringthe capture camera, controllingthe motorfor focusing the capture camera, controllingthe aiming motorsthat allow the capture camera to be oriented, and finally controllingilluminationof the region of interest synchronously with the capture of images by the capture camera. The imagesgenerated by the context camerasand the imagesgenerated by the capture cameraare transmitted for processing to the main processor.

202 105 102 102 The real-time coprocessoris responsible for real-time tasks. It triggers acquisition of the images by the context camera systemat regular intervals, for example at a frequency preferably greater than 7 Hz and equal in the example of embodiment described here to 15 Hz. It controls the movements of the aiming motors, for example via servo-control of position and/or speed. It controls the movements of the focus motor. It triggers image capture by the capture cameraat the right time, and activates illumination of the region of interest synchronously with capture by the capture camera.

201 105 103 104 202 The main processorreceives a stream of images acquired by the context cameras. In these acquired images, also called context images, it detects the presence of the subjectand locates in three dimensions the region of interestin the detection volume. On the basis of these three-dimensional coordinates, the main processor estimates the path of the region of interest and then sends the path to the coprocessor.

202 206 201 The coprocessorservo-controls the position and speed of the aiming motorsso as to reach as quickly as possible the path sent to it by the main processor.

capture capture capture 202 102 When the aiming motors have “locked onto” the ideal path and point in the direction of the determined position of the region of interest at the time t, the coprocessortriggers capture of the image of the region of interest by the capture cameraat the time t, this capture possibly being repeated, in particular periodically, for example in the event of concealment of the subject at the time t.

202 The more often the coprocessorreceives path updates, the closer it will get to the actual path of the region of interest.

The advantage of the synchronization between the main processor and coprocessor is that transmission of commands by the main processor is not subject to real-time constraints, this allowing a multi-task non-real-time operating system to be used. In other alternative embodiments, a single real-time processor handles all of the tasks here shared between the central processor and the real-time coprocessor.

4 FIG. illustrates the software architecture used for path estimation and tracking in one example of embodiment of the invention.

105 The example of embodiment is based on the use of two context camerasused in a stereoscopic mode to allow three-dimensional location of objects detected in the image. It will be noted that other embodiments may use other techniques, alternatively or in addition to stereoscopic imaging, such as time-of-flight cameras, or even a system for achieving three-dimensional vision based on structured light.

201 301 The main processorexecutes a first modulethat is responsible for locating the region of interest, for example the eyes of the subject in the case of recognition of irises within the context images. The recognition and location algorithm used is known per se and is not a subject of this document.

302 201 A second moduleexecuted by the main processoris responsible for computing the three-dimensional position of the region of interest located by the first module. Preferably, the frequency at which the three-dimensional position of the region of interest is computed is less than or equal to the frequency of acquisition by the context camera, for example 15 Hz, the advantage of computing a position at a frequency less than the acquisition frequency of the images being to decrease the load on the processor at times of high demand—the computation frequency may therefore advantageously be varied depending on load. These positions are stored in memory, or at least the latest few are stored in memory. The number of positions stored in memory with a view to estimating the path, which is defined in at least one dimension by a parametric path model combining a periodic temporal component and a linear temporal component, is at least two positions. Preferably, the constituent positions of a complete step of the subject are used in the initialization phase of the extended Kalman filter employed to estimate the path.

303 201 n n n n i eyes periodic walk a periodic temporal component P(t) that is dependent on an amplitude parameter Amplitude, a phase parameter do and an angular-frequency parameter ω. Equivalently to this angular-frequency spectral decomposition, a frequency decomposition is possible linear average path average path a linear temporal component P(t) dependent on a rate-of-change parameter driftand on an intercept parameter P. A third moduleexecuted by the main processoris responsible for estimating the path of the region of interest. The path is estimated as a function of time. To do this, the received context images are time-marked with a timestamp when they are received. These timestamps allow a time index n to be associated with the computed three-dimensional positions. The time index n corresponding to a time t therefore corresponds to the three-dimensional position P=(X, Y, Z). A sequence of indexed positions Pis thus obtained. The path model P(t) comprises:

eyes The parametric path model P(t) of the eyes, in three dimensions, is written:

linear average path average path average path average path 0 P(t)=P(t), which corresponds in this example to the average position of the eyes excluding any oscillation over time, i.e. to the average position of the two eyes as a function of time excluding any oscillation, here in three dimensions, and with P(t)=P+drift*(t−t); periodic 0 walk 0 and P(t)=Amplitude*sin(Φ(t)), which corresponds in this example to the oscillation of the average position of the eyes over time, i.e. to the oscillation of the average position of the two eyes as a function of time, here in three dimensions, and with Φ(t)=Φ+ω*(t−t).

In order to separate the linear component from the periodic component of the instantaneous observable position signal, a band-pass filter, which may be different depending on the dimensions, is applied to the instantaneous observable position signal made up of the computed positions of the eyes over time.

a measurement vector An extended Kalman filter is used to estimate this parametric path model. The extended Kalman filter is based on:

containing two observers, the first observer of which is the instantaneous observable position signal, which consists of the sequence of computed eye positions, and the second observer of which is the periodic observable signal; a state vector

average path average path average path average path eyes  containing five states including an average position P(t) of the region of interest in each dimension, a rate of change P(t) in the average position of the region of interest in each dimension, which therefore corresponds to a representative rate of the moving subject corresponding to the change between two measurement times in the speed of a moving subject (excluding any oscillation), an amplitude of oscillation Amplitude of the region of interest around the average position, a phase Φ(t) of oscillation around the average position and an angular frequency Φ(t) of oscillation of the average position of the region of interest. Thus in this example, the states listed above correspond to the height of the person's eyes (P(t)), to the rate of change in this height (excluding any oscillation (“linear” component)) (P(t)), to the amplitude of the oscillation of the height of the eyes (Amplitude), to the phase of the sinusoid modeling this oscillation around the average position (Φ(t)) and to the rate of change in this phase (Φ(t)), i.e. the angular frequency of oscillation, which in particular is related to the pace of the steps of the person. The amplitude of the oscillation is in particular related to the morphology and gait of the person. Selecting such states makes it possible to increase accuracy by increasing fidelity to parameters describing the person's gait, the first two states relating to the linear part of the prediction and the last three to the periodic part of the prediction. It will be noted that, equivalently, the angular frequency and phase may be replaced by a frequency and phase.For each new subject, the path of their eyes P(t) is estimated by means of a parametric model, through implementation of an extended Kalman filter of the parametric path model in each dimension, this comprising a phase of initialization of the extended Kalman filter, a phase of prediction by the extended Kalman filter and a phase of update of the extended Kalman filter at a sampling frequency. The following illustration will now be used to explain in more detail how this path estimation step is implemented. In the case presented here, only a single path of the eyes is modeled per person because the modeled path is that of a point halfway between the two eyes, the field of the capture camera allowing both eyes to be acquired at the same time, because the camera advantageously comprises two sensors with two lenses that share the same motor; nevertheless, it would be possible to compute the path of each eye and to update the path of each eye in parallel.

102 Advantageously, in the case where the acquired images contain a plurality of regions of interest, for example in the case where a plurality of subjects (each with one region of interest) are detected in the acquired images, the paths of the regions of interest (including in particular filter updates) may be estimated in parallel, the images of each region of interest then possibly being captured sequentially, in particular if there is only one capture camera.

303 304 202 303 303 The path-estimating moduletransmits the successive estimations to a path-tracking moduleexecuted by the real-time coprocessor. The path-tracking module is capable of determining the position of the region of interest, at any time, on the basis of the last path estimation transmitted by the module, so as to ensure continuous optimized pointing. In practice, the frequency of transmission of the successive estimations is generally greater than the frequency of capture of the context images. In this example, the path estimations are transmitted at a frequency of 15 Hz, then in the real-time coprocessor the path-tracking module is responsible for determining the position of the region of interest at a frequency of 1 kHz, on the basis of the last path estimation transmitted by the module, this guaranteeing continuity in the pointing of the capture camera to disengagement from the path, even though the capture frequency is here 15 Hz, or variable. Specifically, if the captured image is suitable for biometric recognition of the subject, another capture is not necessary.

305 102 102 This ability to determine the position of the region of interest at any time makes it possible to control the modulefor controlling the motors used to aim the capture camera. This is how the capture cameracontinuously tracks the region of interest.

306 306 307 102 306 307 100 A modulealso receives the determined position of the region of interest. This moduleis responsible for computing a focus factor that is added to the depth coordinate (Z) of the determined position, i.e. to the distance between the capture camera and the region of interest. The focus factor varies, for example linearly, between a negative minimum and a positive maximum. The distance between the capture camera and the determined position of the region of interest, corrected by the focus factor, allows the modulethat controls the motor used to focus the capture camerato be controlled. The modulesandhave been drawn with dashed lines because they are optional, the devicebeing able to operate without additional focusing, in particular if the depth of field of the lenses of the capture camera is sufficiently wide relative to the error in the position prediction, a few centimeters for example when EDOF sensors are used (EDOF standing for Extended Depth of Focus).

303 i n-1 n-1 n-1 n-1 n n n n n n n n n n-1 n n-1 n n-1 n-1 n-2 Optionally, the third moduleis responsible for an additional estimation of the path of the region of interest. In the phase of prediction by the extended Kalman filter, as soon as the prediction error |Pk/k−1−Pk/k| of the extended Kalman filter, in relation to the computation of the actual position, drops below a convergence threshold, for example one determined statistically or dynamically, the extended Kalman filter is considered to have converged. Advantageously, as long as the extended Kalman filter has not converged the additional estimation is used, in particular during the initialization phase. This additional estimation is typically made through linear approximation of the stored positions. The simplest embodiment merely estimates the coordinates of a straight line passing through the last two stored positions. In this case, a locally rectilinear path is estimated, which may be sufficient. In more complex embodiments, it is possible to compute a polynomial model of the path based on a number of positions greater than two. Such a polynomial model makes it possible to obtain a path estimation that is more accurate than the locally rectilinear path. The path is estimated as a function of time. The sequence of indexed positions Pmakes it possible to compute a speed associated with this position once at least two positions have been stored. This series may have indexes with which there are no positions associated. This may be due to a problem with transmission of context images, or to an inability to recognize the region of interest in certain context images for example. These missing positions must be taken into account when estimating the path, and in particular in the estimation of speed. For example, for a locally rectilinear path estimation based on two stored positions P=(X, Y, Z) and P=(X, Y, Z), it is possible to estimate the speed of the subject at the time t corresponding to time index n: V=(VX, VY, VZ)=(X−X, Y−Y, Z−Z). If position Pis missing, it is possible to use position Pin a similar way, though it is then necessary to divide the obtained speed by two to take into account the missing position. It is then possible to estimate a position at a time T greater than the current time t, from the current position at time t by applying the speed over a time (T−t).

5 FIG. eyes illustrates the principle of path estimation in one example of embodiment of the invention. For the sake of clarity, the description will be focused solely on parametric modeling of the path P(t) only in the y-dimension, which characterizes the height of the eyes of the subject, which is the most subject to periodic variation.

The dynamics of the eyes are modeled using the state equations of the system:

either via Euler discretization, k denoting the time index:

k k k k k k k eyes average path average path average path 0 linear periodic 0 walk 0 with:xthe actual state,uthe input command, here zero given the absence of commands;wthe transition noise, which is a centered Gaussian of covariance matrix Qzmeasurementvthe measurement noise, which is a centered Gaussian of covariance matrix R;or by considering that P(t) is made up of P(t)=P+drift*(t−t) and P(t)=P(t)=Amplitude*sin(Φ(t)), with Φ(t)=Φ+ω*(t−t), so that in the case of variation of the position of the eyes along the y axis the following may be written:

0 walk 0 with height(t) the detected height of the eyes of the subject, hthe initial height of the person's eyes, drift the rate of change in this height (excluding any oscillation (“linear” component)), Amplitude the amplitude of the oscillation of the height of the eyes, Φ(t) the phase of the oscillation, do the initial phase and ωthe rate of change in this phase, i.e. the angular frequency of oscillation, which is in particular related to the pace of the steps of the person. The term “initial” here refers to a time tof the end of the initialization phase.

Applying the extended Kalman filter to the two-observer model, it is then possible to write the state vector X(t) and the measurement vector Z(t) as follows:

303 The modulefor estimating the path of the eyes receives as input the three-dimensional positions of the eyes, which are computed at said computation frequency on the basis of one or more images captured by the context camera system. Each received three-dimensional position is a constituent of the instantaneous observable position signal.

303 a The instantaneous observable position signal is here processed so as to extract only its component along the y-axis:to which is applieda band-pass filter so as to isolate the non-linear part, which is assumed to be periodic, and in particular sinusoidal, of the component along the y-axis of the instantaneous observable position signal. The application frequency of the band-pass filter is for example 15 Hz, and it is in particular equal to the acquisition frequency, or less than the acquisition frequency so as to reduce processor load. Advantageously, the application frequency of the band-pass filter is variable, this making it possible to cope with cases of concealment of the subject, i.e. when the three-dimensional position cannot be computed in certain acquired images. The signaltherefore corresponds to the instantaneous observable position signal along the y-axis, for example as returned directly by two constituent stereoscopic cameras of the context camera system, and the signaltherefore corresponds to the periodic observable signal (centered on zero) extracted from the instantaneous observable position signal, as output from the band-pass filter fed with the y-axis componentof the instantaneous observable position signal. The parameters of the band-pass filter will have been pre-calibrated by spectral analysis of trial data obtained from a sample of subjects, the same sample that made it possible to determine that the average oscillation frequency (in the y-direction) of a walking subject is between 1 and 2 Hz.

303 b 0 an amplitude parameter Amplitude; 0 a phase parameter P; an angular-frequency parameter ω; a rate-of-change parameter drift; 0 0 105 an intercept parameter P;said initialization parameters being determined from all or part of the sequence of computed three-dimensional positions of the region of interest in the detection volume. Preferably, the initialization parameters, forming the initial state vector X(t), are determined from the instantaneous observable position signal in such a way that the part of the sequence covers the captured “half first step” of the subject in the images acquired by the context camera system. To this end, the part of the sequence of positions comprises the positions computed for said region of interest on the basis of images acquired successively by the context camera system, the first image of which is the one in which the region of interest is detected a first time, and the following images of which are those in which the region of interest is detected, until at least two local extrema are detected. The phaseof initialization of the extended Kalman filter of the parametric path model comprises determining an initial state vector X(t) by determining initialization parameters comprising:

the observation function; This makes it possible to write the following equations of the extended Kalman filter:

s  and, for a length of time Tfrom the last estimation, representing the variable sampling period; the transition function (also known as the prediction function):

the transition matrix F and observation matrix H being defined as being Jacobians of f and h, respectively:

303 303 c d In each current time step a phaseof prediction of the state X(t+1) at the subsequent time takes place as does a phaseof updating the extended Kalman filter with respect to the measurementof the current time step) and the prediction made in the previous time step.

In other words, during the prediction phase (also called the convergence phase) the state estimated at the previous time is used to produce an estimation of the current state:

k k with P the error covariance matrix, i.e. here the covariance matrix of the measurement noise because there is considered to be no process noise, uzero because the motion of the subject is not controlled, and Fa matrix that relates the previous state k−1 to the current state k.

It is written:

measure_bpf|measure measure|measure_bpf measure measure_bpf 100 with, considering independent noises: P=0 and P=0, and the values of Pand Pdetermined experimentally and statistically for the devicein question.

s drift angular frequency For a path estimation at constant frequency 1/T, the transition noise depends solely on the noise at said frequency (qor q)

The initial covariance (uncertainty) matrix Q of the transition noise of the model

drift the variance of the speed of the people in question: qrepresenting the change from one time to the next in the speed of a person (linear part of the prediction); angular frequency the variance qof the period of oscillation; and a the variance qof the amplitude of oscillation;all three for example being preset empirically. In particular, the initial coefficients of the covariance matrix Q of the extended Kalman filter are obtained through successive iterations minimizing the error in the model on paths of the trials database. is calibrated based on:

303 303 303 c c d. In the phaseof prediction by the extended Kalman filter, provided that the prediction error |Pk/k−1−Pk/k| of the extended Kalman filter, in relation to the computation of the actual position, is greater that a convergence threshold, the extended Kalman filter is not considered to have converged. This check also ensures at each time step that the filter has not diverged. The convergence threshold is advantageously determined statistically or dynamically—it is for example of the order of one meter, or less, and in particular 0.2 m. Independently of the convergence of the filter, the prediction phasecontinues with the update phase

In the update phase, the observations of the current time are used to correct the predicted state with the aim of obtaining a more accurate estimation:

The matrix P is updated in each computation cycle corresponding to a new computation of position and gives the level of confidence in the path. Preferably, in case of loss of the region of interest (concealment), the filter is not updated and when the region of interest is again detected on a subsequent acquisition, the filter is updated with the sampling period corresponding to the length of time between the time when the region of interest was acquired (and de facto its position computed) last and the time when the region of interest was again acquired.

measure position If the extended Kalman filter has converged, the estimation of the path of the region of interest at said future time tracking is obtained by computing the tangent to the path modeled on the basis of the acquisitions made up to the time of the last acquisition tused to update the extended Kalman filter (i.e. including the region of interest), in particular by writing the modeled path in the form of a position pos and speed spd:

tracking 202 If tis selected to be the estimated time of receipt of said path estimated by the real-time coprocessor, this makes it possible to compensate for the latency between the real times and the last acquisition used to update the extended Kalman filter.

tracking 304 The estimation of the path in the y-dimension of the eyes at a given time tof receipt by the path-tracking moduleis obtained by computing the tangent to the modeled path and is written:

The path in the other two dimensions is for example obtained using the model described in connection with the additional path estimation.

capture capture capture 305 307 Thus, at the current time, the camera is commanded to point toward the coordinates of the position pos (t) of the region of interest estimated for the future time tby applying the preceding equation to the time t, this making it possible to compensate for the response time of the pointing systemand/or focusing systemof the capture camera.

To obtain a reliable path estimation, it is necessary to sample the position computation at a sufficient frequency. This frequency will depend on the speed of movement of the subject and therefore depend on the envisioned application. For example, in the case of recognition of the irises of a walking person, a frequency of 15 context images and therefore of 15 computed position estimations per second has proved sufficient, the same going for the path estimation.

6 FIG. 106 106 601 605 604 607 602 603 106 shows one example of a data-processing devicefor implementing one or more embodiments of the invention. The data-processing devicetypically comprises one or more central processing units (CPUs)and/or one or more graphics processing units (GPUs), a physical communication module (NET), one or more physical input/output modulesfor interchanging data with external devices (such as the context camera system and the capture camera), a transient storage mediumsuch as a random access memory (RAM), a non-transient recording medium(FLASH) and communication buses (not shown) for transferring data between the internal components of the data-processing device.

106 301 302 303 304 305 306 307 106 The data-processing devicemay be used to execute one or more program modules,,,,,,comprising instructions that, when the program module or modules are executed, cause the data-processing deviceto carry out the method according to the invention. The program module or modules may be written in any, compiled or interpreted, programming language. They may form part of a software solution, i.e. of a collection of executable instructions, of codes, of scripts or the like and/or of databases.

106 601 a central processing unit (CPU), such as a microprocessor, and including in particular an internal clock; 602 602 a transient memory, for storing the executable code of the method for carrying out the invention and registers configured to record variables and parameters required to implement the method according to embodiments of the invention; the memory capacity of the device is preferably supplemented by an optional random-access memoryconnected to an extension port, for example; 603 106 603 a non-transient memoryfor storing the computer programs and calibration data needed to implement embodiments of the invention; the stored computer programs in particular comprise a computer program comprising instructions configured to implement all or some of the steps of the method according to the invention when said program is executed on the processing device, said non-transient memorythen being one example of a (removable or irremovable) non-transient information storage means; 604 604 604 601 a communication modulecomprising a network interfaceconnected to a communication network over which digital data to be processed are transmitted or received; the network interfacemay be a single network interface, or be made up of a set of different network interfaces (for example wired and wireless interfaces or different types of wired or wireless interfaces). Data packets are sent to the network interface for transmission or are read from the network interface when received under control of the software application executed in the processor; 605 a user interface (HMI), in particular one comprising a graphics processor, for receiving inputs from a user or for displaying information to a user, in particular (visual and/or vocal) guidance information; 607 an input/output modulefor receiving/sending data from/to external peripherals such as a hard disk, a removable storage media, inter alia. The data-processing devicecomprises the following elements, connected to each other via a communication bus:

603 604 106 603 The executable code may be stored in the non-transient memory, for example a flash memory or a read-only memory, or on a removable digital medium such as, for example, a disk. According to one variant, the executable code of the programs may be received by means of a communication network, via the network interface, in order to be stored in one of the storage means of the data-processing device, such as the memory, before being executed.

601 603 601 602 601 The central processing unitis configured to control and direct execution of the instructions or of segments of software code of the program or programs according to one of the embodiments of the invention, which instructions are stored in one of the aforementioned storage means, such as the non-transient memory. After being turned on, the CPUis capable of executing instructions from the transient RAM memorythat relate to a software application. Such software, when executed by the processor, allows the method according to the invention to be executed.

In one embodiment, the device is a programmable device that uses software to implement the invention. As a variant, the present invention may be implemented in hardware form (for example, in the form of an application-specific integrated circuit (ASIC) or in the form of a field-programmable gate array (FPGA)).

106 1 604 According to one embodiment, the data-processing deviceis solely hosted locally, or, as a variant, is external to the terminal, or even distributed and comprises multiple processing sub-units, which in particular are at least partly external and communicate with one another via the network interface. Similarly, depending in particular on the nature of the terminal, all or part of the memory may be physically remote, hosted for example on a remote server.

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

January 28, 2026

Publication Date

September 10, 2026

Inventors

Ngoc-Son-Dorian NGUYEN
Yannick LOITIERE

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “METHOD AND DEVICE FOR CAPTURING IMAGES OF A MOVING SUBJECT” (US-20260270561-A1). https://patentable.app/patents/US-20260270561-A1

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.