Patentable/Patents/US-20260229023-A1
US-20260229023-A1

Learning Method and Learning Apparatus

PublishedAugust 6, 2026
Assigneenot available in USPTO data we have
InventorsTakahiro AOKI
Technical Abstract

A processing unit converts a first learning image, which is for machine learning of a neural network configured to calculate authentication data for biometric authentication from an input biometric image, into first three-dimensional data by using a first conversion expression based on internal parameters of an imaging apparatus for capturing the biometric image; generates second three-dimensional data by varying a posture of a biological body part included in the first three-dimensional data; generates a second learning image by converting the second three-dimensional data by using a second conversion expression that performs inverse conversion of conversion performed by using the first conversion expression; and executes machine learning of the neural network by using at least the first learning image and the second learning image.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

converting a first learning image, which is for machine learning of a neural network configured to calculate authentication data for biometric authentication from an input biometric image, into first three-dimensional data by using a first conversion expression based on an internal parameter of an imaging apparatus for capturing the biometric image; generating second three-dimensional data by varying a posture of a biological body part included in the first three-dimensional data; generating a second learning image by converting the second three-dimensional data by using a second conversion expression that performs inverse conversion of conversion performed by using the first conversion expression; and executing machine learning of the neural network by using at least the first learning image and the second learning image. . A non-transitory computer-readable storage medium storing a computer program that causes a computer to execute a process comprising:

2

claim 1 . The non-transitory computer-readable storage medium according to, wherein the first conversion expression and the second conversion expression are generated based on the internal parameter and a predetermined distance from the imaging apparatus to the biological body part to be imaged.

3

claim 1 . The non-transitory computer-readable storage medium according to, wherein an amount of variation of the posture of the biological body part included in the first three-dimensional data is determined probabilistically.

4

claim 1 . The non-transitory computer-readable storage medium according to, wherein the generating of the second three-dimensional data includes, varying, by probabilistically setting a rotation angle and a translation amount in a three-dimensional space, at least one of the posture and a position of the biological body part included in the first three-dimensional data.

5

claim 4 . The non-transitory computer-readable storage medium according to, wherein the generating of the second three-dimensional data includes setting, as a setting probability of the translation amount in a z-axis direction, a probability of translation in a negative direction higher than a probability of translation in a positive direction.

6

claim 1 . The non-transitory computer-readable storage medium according to, wherein the converting of the first learning image into the first three-dimensional data includes executing normalization processing on the first learning image to correct the posture of the biological body part on the first learning image to a predetermined posture, thereby generating a normalized image, and converting the normalized image into the first three-dimensional data by using the first conversion expression.

7

claim 1 . The non-transitory computer-readable storage medium according to, wherein the executing of the machine learning of the neural network includes executing machine learning of a spatial transformer network (STN) together with the executing of the machine learning of the neural network, the spatial transformer network being disposed in front of the neural network and configured to correct the posture of the biological body part on the input biological image.

8

claim 1 identification numbers for identifying imaging target persons are associated with respective ones of a plurality of learning images including the first learning image, the executing of the machine learning of the neural network includes generating a classifier for classifying the imaging target persons by executing the machine learning using the plurality of learning images and the second learning image as training data and using the identification numbers as ground truth data, and using, as an identification number associated with the second learning image, the same identification number associated with the first learning image, and the neural network, in biometric authentication processing, calculates, as the authentication data, a feature vector output from an intermediate layer included in the neural network upon receiving the biometric image. . The non-transitory computer-readable storage medium according to, wherein:

9

converting, by a processor, a first learning image, which is for machine learning of a neural network configured to calculate authentication data for biometric authentication from an input biometric image, into first three-dimensional data by using a first conversion expression based on an internal parameter of an imaging apparatus for capturing the biometric image; generating, by the processor, second three-dimensional data by varying a posture of a biological body part included in the first three-dimensional data; generating, by the processor, a second learning image by converting the second three-dimensional data by using a second conversion expression that performs inverse conversion of conversion performed by using the first conversion expression; and executing, by the processor, machine learning of the neural network by using at least the first learning image and the second learning image. . A learning method comprising:

10

a memory; and convert a first learning image, which is for machine learning of a neural network configured to calculate authentication data for biometric authentication from an input biometric image, into first three-dimensional data by using a first conversion expression based on an internal parameter of an imaging apparatus for capturing the biometric image, generate second three-dimensional data by varying a posture of a biological body part included in the first three-dimensional data, generate a second learning image by converting the second three-dimensional data by using a second conversion expression that performs inverse conversion of conversion performed by using the first conversion expression, and execute machine learning of the neural network by using at least the first learning image and the second learning image. a processor coupled to the memory and the processor configured to: . A learning apparatus comprising:

Detailed Description

Complete technical specification and implementation details from the patent document.

This application is a continuation application of International Application PCT/JP2023/036622 filed on Oct. 6, 2023, which designated the U.S., the entire contents of which are incorporated herein by reference.

The embodiments discussed herein relate to a learning method and a learning apparatus.

In recent years, biometric authentication using various biological body parts has been utilized. For example, in fingerprint authentication, authentication is often performed by bringing a biological body part (a finger) into contact with a biometric sensor. On the other hand, in palm vein authentication and iris authentication, authentication is performed in a non-contact manner.

In addition, in biometric authentication processing, there are cases where a neural network generated by machine learning is utilized. For example, it has been considered to utilize a neural network in processing for calculating registration authentication data and matching authentication data from a captured biometric image.

International Publication Pamphlet No. WO 2014/039732 U.S. Patent Application Publication NO. 2022/0366717 As a related technique of biometric authentication, a biometric authentication system has been proposed in which a neural network serves as a local classifier for classifying a palm region as a region of interest. In addition, as a related technique of neural networks, a data labeling method has been proposed in which, in order to train a neural network that performs hand tracking, three-dimensional position coordinates of bone points are acquired based on depth images collected by a depth camera, and joint information of the bone points is labeled based on the three-dimensional position coordinates and two-dimensional position coordinates on the depth images. See, for example, the following literatures.

In one aspect, there is provided a non-transitory computer-readable storage medium storing a computer program that causes a computer to execute a process including: converting a first learning image, which is for machine learning of a neural network configured to calculate authentication data for biometric authentication from an input biometric image, into first three-dimensional data by using a first conversion expression based on an internal parameter of an imaging apparatus for capturing the biometric image; generating second three-dimensional data by varying a posture of a biological body part included in the first three-dimensional data; generating a second learning image by converting the second three-dimensional data by using a second conversion expression that performs inverse conversion of conversion performed by using the first conversion expression; and executing machine learning of the neural network by using at least the first learning image and the second learning image.

The object and advantages of the invention will be realized and attained by means of the elements and combinations particularly pointed out in the claims.

It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are not restrictive of the invention.

In non-contact biometric authentication, there are cases where a posture of a biological body part with respect to a biometric sensor is not constant. Variations in the posture of the biological body part become a major factor causing authentication errors.

Hereinafter, embodiments of the present invention will be described below with reference to the drawings.

1 FIG. 1 FIG. 10 1 illustrates an example of a configuration and processing of a learning apparatus according to a first embodiment. A learning apparatusillustrated inis an apparatus that generates, by machine learning, a neural networkwhich h is used at the time of biometric authentication processing. As the biometric authentication, non-contact biometric authentication is applied in which biological information is read in a state where a biological body part does not contact a biometric sensor, such as palm vein authentication.

10 11 12 11 10 13 11 13 12 10 12 The learning apparatusincludes a storing unitand a processing unit. The storing unitis a storage area secured in a storage device included in the learning apparatus. Internal parametersof a biometric sensor are stored in the storing unit. The internal parametersinclude, for example, information such as a focal length and image center coordinates. The processing unitis, for example, a processor included in the learning apparatus. The following processing performed by the processing unitis implemented, for example, by the processor executing a program.

1 1 2 1 2 2 1 2 11 n n When a biometric image is input, the neural networkcalculates authentication data for biometric authentication (registration data and matching data). In addition, training of the neural networkis executed using a plurality of learning images prepared in advance. The plurality of learning images is biometric images obtained by capturing a biological body part. In the present embodiment, as an example, it is assumed that n learning images-to-are prepared in advance. The learning images-to-may be stored in the storing unit.

12 2 1 2 2 1 12 2 1 3 13 n a The processing unitexecutes training data augmentation processing (augmentation) for increasing the number of learning images by selecting one or more learning images from the learning images-to-. Here, as an example, it is assumed that the learning image-is selected. The processing unitconverts the selected learning image-into three-dimensional databy a first conversion formula based on the internal parameters.

12 3 3 4 3 4 4 4 a b a 1 FIG. Next, the processing unitconverts the three-dimensional datainto three-dimensional databy varying a posture of a biological body partincluded in the three-dimensional data. In this processing, a three-dimensional posture variation is added to the biological body part. In the example of, the biological body partis a palm, and the posture of the biological body partis varied so that a state in which a surface of the palm is substantially parallel to an x-y plane in a three-dimensional coordinate system becomes a state in which an inclination occurs with respect to the x-y plane.

12 2 1 3 2 1 4 2 1 a b a Next, the processing unitgenerates a learning image-by converting the three-dimensional datainto two-dimensional data by a second conversion formula that performs inverse conversion of the first conversion formula. In this manner, the learning image-, in which the biological body partis imaged in a state where the posture is varied, is generated from the learning image-.

1 12 1 2 1 2 1 1 2 1 2 2 1 a n a During training of the neural network, the processing unitexecutes machine learning of the neural networkby using at least the learning image-and the learning image-. In practice, machine learning of the neural networkis executed using the learning images-to-together with a plurality of newly generated learning images such as the learning image-generated from the plurality of original learning images.

2 1 2 1 4 4 4 4 1 1 4 1 4 n As a result, compared with a case where training is performed using only the learning images-to-, the neural networkcapable of calculating authentication data that is robust to variations in the posture of the biological body partis generated. In particular, in the above processing, since it becomes possible to vary the biological body partin three-dimensional space in a state where the biological body partis converted into three-dimensional data, learning images that accurately reproduce posture variations of the biological body partthat may actually occur are additionally generated. Then, by training the neural networkusing the additionally generated learning images together with the original learning images, the neural networkcapable of reliably calculating authentication data that is robust to variations in the posture of the biological body partis generated. By using such a neural networkfor authentication processing, highly accurate authentication processing robust to variations in the posture of the biological body partis executed.

Here, as the above second conversion formula, a camera matrix that converts three-dimensional coordinates in a three-dimensional space into two-dimensional coordinates may be used. In non-contact biometric authentication, there are cases where an appropriate distance exists between a biometric sensor and a biological body part at the time when a biometric image is captured. In addition, there are many cases where the biometric sensor and the biological body part are substantially parallel at the time when the biometric image is captured. In such cases, among parameters of the camera matrix, a parameter indicating a distance of the biological body part in a z-axis direction is set to a fixed value indicating the appropriate distance, and a rotation parameter is set to zero. As a result, by inverse conversion using the camera matrix, the learning image is converted into three-dimensional data, and by using the camera matrix, the three-dimensional data after posture variation is converted into two-dimensional data.

Next, a system in which vein authentication using palm veins is executed as an example of biometric authentication will be described.

2 FIG. 2 FIG. 100 200 illustrates an example of a configuration of a biometric authentication system according to a second embodiment. As illustrated in, the biometric authentication system includes a learning apparatusand an authentication apparatus.

200 201 201 The authentication apparatusincludes a biometric sensorfor acquiring biometric information of a user. In the present embodiment, information of palm veins is used as the biometric information. In this case, the biometric sensorincludes, for example, a light-emitting unit that emits near-infrared light onto a palm, and an imaging unit that captures a biometric image by receiving reflected near-infrared light from the palm.

200 200 200 The authentication apparatusgenerates authentication data used for authentication by performing image processing on a captured biometric image. The authentication apparatusgenerates registration authentication data (registration data) from the biometric information of a user to be registered, and registers the generated registration data in a database. In addition, the authentication apparatusgenerates matching authentication data (matching data) from the biometric information of a user to be authenticated, and outputs an authentication result by matching the generated matching data with the registration data registered in the database.

200 100 In the authentication apparatus, authentication data is generated by using a convolutional neural network (CNN). The learning apparatusgenerates the above CNN by executing machine learning using biometric images for training (learning images) captured in advance.

200 200 By training the CNN, weight coefficients between nodes in the CNN are calculated, and a weight dataset including the calculated weight coefficients is generated. The generated weight dataset is delivered to the authentication apparatus, and the authentication apparatusbecomes capable of generating authentication data by setting the weight coefficients included in the weight dataset in the CNN.

100 200 100 200 200 The learning apparatusis connected to the authentication apparatus, for example, via a network. In this case, the learning apparatusdelivers the generated weight dataset to the authentication apparatusvia the network. Alternatively, the generated weight dataset may be delivered to the authentication apparatusvia a portable recording medium.

3 FIG. 3 FIG. 3 FIG. 100 100 101 102 103 104 105 106 107 illustrates an example of a hardware configuration of the learning apparatus. The learning apparatusis implemented, for example, as a computer such as that illustrated in. As illustrated in, the learning apparatusincludes a processor, a random access memory (RAM), a hard disk drive (HDD), a graphics processing unit (GPU), an input interface, a reading device, and a communication interface.

101 100 101 101 101 12 100 101 100 101 1 FIG. The processorintegrally controls the entire learning apparatus. The processoris, for example, a central processing unit (CPU), a micro processing unit (MPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a programmable logic device (PLD). The processormay also be a combination of two or more of a CPU, an MPU, a DSP, an ASIC, and a PLD. The processoris an example of the processing unitillustrated in. The learning apparatusmay include a plurality of processors. Different processors may be used to execute different processes among a plurality of processes performed by the learning apparatus. The processormay be referred to as processor circuitry.

102 100 101 102 101 102 The RAMis used as a main storage device of the learning apparatus. At least a part of an operating system (OS) program and an application program to be executed by the processoris temporarily stored in the RAM. Various data needed for processing by the processorare also stored in the RAM.

103 100 103 The HDDis used as an auxiliary storage device of the learning apparatus. An OS program, application programs, and various data are stored in the HDD. As the auxiliary storage device, another type of nonvolatile storage device such as a solid state drive (SSD) may be used.

104 104 104 104 101 104 a a a A display deviceis connected to the GPU. The GPUcauses an image to be displayed on the display devicein accordance with an instruction from the processor. Examples of the display deviceinclude a liquid crystal display and an organic electroluminescence (EL) display.

105 105 105 105 101 105 a a a An input deviceis connected to the input interface. The input interfacetransmits a signal output from the input deviceto the processor. Examples of the input deviceinclude a keyboard and a pointing device. Examples of the pointing device include a mouse, a touch panel, a tablet, a touch pad, and a trackball.

106 106 106 106 101 106 a a a A portable recording mediumis attachable to and detachable from the reading device. The reading devicereads data recorded on the portable recording mediumand transmits the data to the processor. Examples of the portable recording mediuminclude an optical disc and a semiconductor memory.

107 200 The communication interfacetransmits and receives data to and from another apparatus such as the authentication apparatusvia a network.

100 200 By such a hardware configuration, processing functions of the learning apparatusare implemented. The authentication apparatusmay also be implemented as a computer including a processor and a memory.

4 FIG. 200 50 50 100 50 illustrates a convolutional neural network (CNN) for explaining generation of authentication data. As described above, the authentication apparatusgenerates authentication data by using a CNN. The CNNis trained by the learning apparatus. The training of the CNNis performed as follows.

100 50 100 50 1 n 1 n 1 n 1 n 1 n 1 n 1 n 1 n 1 n In the learning apparatus, n learning images Ito Iare prepared for training the CNN. The learning images Ito Iare biometric images obtained by capturing palms of different persons, respectively. Identification information IDto IDfor identifying the persons who are imaging targets is associated with the respective learning images Ito I. The learning apparatustrains a classifier having a number of categories equal to n, whereby an input image is classified into the n persons corresponding to the learning images Ito I, by using the learning images Ito Ias training data and IDto IDas ground truth data. As a result, the CNNthat calculates probability values Pto Pcorresponding to the respective IDto IDwhen an image is input is generated.

200 50 50 201 50 50 51 in 1 2 m 1 2 m in The authentication apparatusgenerates authentication data by using the CNNgenerated in this manner. As the authentication data, a feature vector composed of a plurality of features is calculated. The CNNincludes an input layer, a plurality of intermediate layers (hidden layers), and an output layer. When an input image Ifrom the biometric sensoris input to the CNN, the CNNoutputs, from a predetermined intermediate layer, an m-dimensional feature vector (V, V, . . . , V). The output feature vector (V, V, . . . , V) becomes information indicating features of the input image I.

1 2 m The feature vector (V, V, . . . , V) may be output from two or more predetermined intermediate layers. As numerical examples of n and m, n=10,000 and m=512 are applicable.

5 FIG. 201 illustrates normalization processing of an input image. In a biometric image used at the time of authentication, it is desirable that an authentication region (a palm in the present embodiment) be imaged in a fixed posture. However, in practice, a user does not always place the authentication region at a correct angle with respect to the biometric sensor. Therefore, captured biometric images often include posture variations of the authentication region. Such posture variations are a major cause of biometric authentication errors, and reducing posture variations is an important factor for improving authentication accuracy.

5 FIG. 200 201 200 50 Accordingly, as illustrated in, the authentication apparatusperforms “normalization processing” on an input image obtained from the biometric sensor, in which a posture of a palm in the image is corrected to a proper posture. The authentication apparatusthen inputs a normalized image obtained by the normalization processing to the CNN. In the normalization processing, rotation and translation are performed so that the posture and position of the palm in the image become a predetermined state. For example, in rotation processing, contour lines of the palm are detected from the input image, a rotation angle is calculated based on an angle of a line extending substantially vertically among the detected contour lines, and the image is rotated by the calculated rotation angle.

50 50 By inputting the normalized image after such processing to the CNN, a feature vector for which authentication errors are unlikely to occur is calculated. However, even if such normalization processing is performed, the posture of the palm is not always corrected accurately. Therefore, the CNNthat processes the normalized image is also expected to exhibit robustness against posture variations.

50 50 50 For this reason, in the present embodiment, augmentation processing is performed on learning images used for training the CNN. In the augmentation processing, a new learning image (an augmented image) is generated by intentionally varying the posture of the palm in an original learning image, and the CNNis trained using the augmented image together with the original learning image. For example, by using as a learning image an augmented image obtained by rotating the palm in the original learning image, the CNNthat is robust to rotation is generated.

As augmentation processing of learning images in a CNN for image processing, for example, the following two methods are conceivable. As a first method, there is a method of using affine transformation. In this method, two-dimensional coordinates of an image are transformed by calculation using a matrix that performs scaling, rotation, and translation. As a second method, there is a method of using homography transformation (projective transformation). In this method, four pairs of points before movement and after movement are selected from the image. Since the number of independent parameters in homography transformation is eight, a transformation matrix (a homography matrix) is uniquely determined by selecting four pairs of movement points, and the two-dimensional coordinates of the image are transformed by the determined matrix.

50 50 However, in each of the first and second methods, three-dimensional variations of a subject are expressed only in a pseudo three-dimensional manner. Therefore, even if augmentation processing is performed on a learning image in which a palm is imaged by either of the above methods, it is not possible to generate an augmented image that accurately expresses three-dimensional posture changes of a palm that actually occur. Accordingly, when the CNNis trained using augmented images generated by the above methods, it is not always possible to generate the CNNthat enables high-performance authentication processing that is robust to posture variations of the palm.

w w w Here, processing for converting three-dimensional coordinates (X, Y, Z) in a global coordinate system into two-dimensional coordinates (u, v) on an image is executed using the following Expression (1) with a camera matrix P.

x y x y 11 12 32 33 1 2 3 In Expression (1), s is a variable indicating a magnification factor. The camera matrix P is expressed as a product of a 3×3 internal parameter matrix including f, f, c, and c, and a 3×4 external parameter matrix including r, r, . . . , r, r, and t, t, and t.

x y x y x y x y 201 The internal parameter matrix is a matrix that converts three-dimensional coordinates in a camera coordinate system into two-dimensional coordinates on an image. fand findicate focal lengths in an x-direction and a y-direction, respectively, and (c, c) indicates image center coordinates. f, f, c, and cof the internal parameter matrix indicate internal parameters of the camera and are determined in advance according to specifications of an image sensor of the biometric sensor.

11 12 32 33 1 2 3 The external parameter matrix is a matrix that converts three-dimensional coordinates in the global coordinate system into three-dimensional coordinates in the camera coordinate system. Among the external parameter matrix, a 3×3 matrix including r, r, . . . , r, and ris a rotation matrix indicating rotation, and a 3×1 matrix including t, t, and tis a translation vector indicating translation.

In the present embodiment, by using such a camera matrix P, an augmented image including a palm to which three-dimensional posture variation is accurately applied is generated.

6 FIG. illustrates a procedure for generating an augmented image.

100 100 100 100 The learning apparatusperforms normalization processing on an original learning image (original image) to generate a normalized image. Next, the learning apparatusperforms inverse conversion using the camera matrix P on the generated normalized image, thereby converting two-dimensional data of the normalized image into three-dimensional data. The learning apparatusthen applies posture variation to the converted three-dimensional data. As a result, three-dimensional variation is added to a subject of the original image in a state where the subject is placed in a three-dimensional space. Thereafter, the learning apparatusconverts the three-dimensional data after the variation into two-dimensional data by using the camera matrix P. In this manner, an augmented image is generated.

1 12 32 33 1 2 3 In general, in conversion from a two-dimensional image to a three-dimensional image, height in a z-axis direction becomes indeterminate. In addition, the external parameter matrix including r, r, . . . , r, r, t, t, and tis not determinable from a single captured image.

0 0 0 201 201 201 201 201 201 However, in non-contact biometric authentication such as palm vein authentication, an appropriate imaging distance (Z) from the biometric sensorat the time of reading biometric information is preset for each biometric sensor. Accordingly, a captured image obtained by the biometric sensormay be approximated as an image in which a distance from the biometric sensoris Zand a surface of a biological body part to be imaged is parallel to an imaging plane of the biometric sensor. In particular, a normalized image to which normalization processing is applied may be assumed to be an image in which the distance from the biometric sensoris closer to Zand the surface of the biological body part is substantially parallel to the imaging plane.

0 1 2 3 0 11 12 32 33 201 100 100 6 FIG. From this, the height in the z-axis direction is set to the above Zindicating the appropriate distance from the biometric sensorto the palm. The matrix elements t, t, and tof the external parameter matrix are determined by the fixed value Z, and the matrix elements r, r, . . . , r, and rof the external parameter matrix become zero on the assumption that the surface of the palm is parallel to the imaging plane. Therefore, the learning apparatusperforms conversion into three-dimensional data from a single learning image as illustrated in, and three-dimensional posture variation of the palm is added by applying a matrix for posture variation to the three-dimensional data. The learning apparatusthen generates an augmented image including a palm to which three-dimensional posture variation is accurately added, by converting the three-dimensional data after the variation into two-dimensional data.

x y x y 201 The units of f, f, c, and cdescribed above are millimeters. Therefore, in order to obtain coordinates in pixel units on an image, an actual pixel size (mm) of an imaging element in the biometric sensoris needed.

4 6 FIGS.to 50 100 200 50 In, for ease of explanation, it is assumed that a biometric image is input directly to the CNNin the learning apparatusand the authentication apparatus. In practice, however, feature extraction processing for detecting a vein pattern is executed on a normalized biometric image, and a feature image after the processing is input to the CNN.

Furthermore, generation of augmented images as described above is not limited to a case where palm veins are used as biometric information. The present technique is also applicable to other types of biometric information in which an appropriate distance from a biometric sensor to an authentication target site of a biological body is predetermined. Examples include iris authentication, non-contact fingerprint authentication in which a distance from a biometric sensor to a finger is kept constant by a physical guide or the like, and face authentication in which a position at which a person stands is specified or a size of a face imaged by the biometric sensor is guided to be constant.

100 Next, details of processing of the learning apparatuswill be described.

7 FIG. 100 110 121 122 123 124 125 illustrates an example of a configuration of processing functions included in the learning apparatus. The learning apparatusincludes a storing unit, a training control unit, a normalization processing unit, a feature extracting unit, a data augmentation processing unit, and a CNN training unit.

110 100 102 103 111 112 113 110 The storing unitis a storage area secured in a storage device included in the learning apparatus, such as the RAMand the HDD. A learning image database (DB), a sensor dataset, and a weight datasetare stored in the storing unit.

50 111 201 200 112 50 113 A plurality of learning images used for learning of the CNNis registered in advance in the learning image database. Identification information (ID) for identifying a person who is an imaging target is added to each learning image. Information indicating specifications of the biometric sensorused in the authentication apparatusis registered in the sensor dataset. Weight coefficients between nodes included in the CNNgenerated by training are registered in the weight dataset.

121 122 123 124 125 101 The processing of the training control unit, the normalization processing unit, the feature extracting unit, the data augmentation processing unit, and the CNN training unitis implemented, for example, by the processorexecuting a predetermined application program.

121 50 122 111 123 The training control unitcontrols training processing of the CNNin an integrated manner. The normalization processing unitperforms normalization processing on a learning image acquired from the learning image database, and outputs a normalized image. The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the normalized image, and outputs a feature image.

124 124 125 50 50 125 110 113 The data augmentation processing unitexecutes augmentation processing for augmenting an image (feature image) for training. The data augmentation processing unitgenerates an augmented image in which a posture of a palm appearing in the feature image is varied three-dimensionally. The CNN training unitexecutes training processing of the CNNby using the original feature image before augmentation processing and the augmented image, and determines weight coefficients between nodes included in the CNN. The CNN training unitstores, in the storing unit, a dataset of the determined weight coefficients as the weight dataset.

8 FIG. 112 201 201 201 x y x y 0 0 illustrates an example of data included in the sensor dataset. In the sensor dataset, focal lengths fand fin x- and y-directions of an imaging element of the biometric sensor, image center coordinates (c, c) of the imaging element, a cell size elen of the imaging element, and a biological body height Zare registered. The biological body height Zindicates a height (appropriate distance) of a biological body part with respect to the biometric sensor. These parameters are determined in accordance with specifications of the biometric sensor. Units of these parameters are all millimeters.

100 Next, processing of the learning apparatuswill be described with reference to flowcharts.

9 FIG. 11 121 112 110 112 [Step S] The training control unitreads the sensor datasetfrom the storing unit, and calculates parameters of the camera matrix P on the basis of the sensor dataset. 12 121 50 [Step S] The training control unitinitializes weight coefficients of the CNN. For example, initial values of the weight coefficients are set by random values. 13 121 111 111 [Step S] The training control unitextracts a predetermined number (for example, approximately sixteen) of learning images from the learning image database. In this processing, the predetermined number of learning images may be randomly extracted from among the learning images registered in the learning image database, or the learning images may be sequentially extracted in groups of the predetermined number. 14 50 [Step S] Training data augmentation processing is executed. In this processing, a feature image for training of the CNNis generated on the basis of each of the predetermined number of learning images that have been extracted. 15 125 50 50 [Step S] The CNN training unitexecutes training processing of the CNNby using each generated feature image, and updates the weight coefficients of the CNN. In this processing, each feature image is used as training data, and machine learning is performed in which ID of a person associated with the learning image serving as the source of the feature image is used as ground truth data. is a flowchart illustrating an example of training processing.

125 50 50 125 50 16 121 13 15 13 17 [Step S] The training control unitdetermines whether to terminate the training processing. For example, when the processing of steps Sthrough Shas been executed a predetermined number of times, it is determined that the training processing is to be terminated. When the training processing is continued, the processing proceeds to step S, and when the training processing is terminated, the processing proceeds to step S. 17 125 110 113 50 [Step S] The CNN training unitstores, in the storing unitas the weight dataset, the weight coefficients that are set in the CNN. The CNN training unitcalculates a training loss when each feature image is input to the CNN. The training loss is a value indicating an error between a value calculated by the CNNupon input of the feature image and the ground truth value. The CNN training unitupdates the weight coefficients of the CNNby performing optimization calculation using an optimization algorithm on the basis of the calculated training loss. As the optimization algorithm, for example, stochastic gradient descent (SGD) or adaptive moment estimation (Adam) is used.

10 FIG. 9 FIG. 10 FIG. 14 13 21 122 [Step S] The normalization processing unitexecutes normalization processing on the learning image. As a result, the posture of the palm in the learning image is corrected such that the posture becomes approximately appropriate, and a normalized image is generated. Note that, in the normalization processing, processing for correcting the palm from a closed state to an opened state may also be executed. 22 123 [Step S] The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the normalized image, and generates a feature image. In this feature extraction processing, for example, vein-enhancement processing such as a band-pass filter is applied to the normalized image, and noise-reduction processing is further applied. 23 124 11 9 FIG. [Step S] The data augmentation processing unitconverts two-dimensional data of the feature image into three-dimensional data, by using a matrix for inverse conversion of the conversion performed by the camera matrix P in which the parameters calculated in step Sofare set. As a result, coordinates of the vein pattern in the feature image are converted into coordinates in a three-dimensional space. 24 124 [Step S] The data augmentation processing unitprobabilistically selects variation parameters for varying a posture and a position of the vein pattern in the three-dimensional space. Specifically, for each variation parameter, a predetermined probability is assigned to each value settable for that parameter, and one of the values is selected according to the assigned probabilities. As one example, the same probability is set for each of the settable values, and one of those values is selected with the same probability (that is, randomly). is a flowchart illustrating an example of the training data augmentation processing. In step Sof, the processing inis executed for each of the learning images extracted in step S.

10 FIG. In the processing of, as an example, not only the posture of the vein pattern is varied but also the position of the vein pattern is varied. As variation parameters for posture variation, rotation angles about each of an x-axis, a y-axis, and a z-axis are used. As variation parameters for position variation, translation amounts in each of an x-axis direction, a y-axis direction, and a z-axis direction are used.

10 FIG. In, as one example, predetermined discrete values are used as the rotation angles and translation amounts. As the discrete values corresponding to each variation parameter, zero, one or more positive values, and one or more negative values are included. Then, one of the discrete values is randomly selected for each variation parameter. Note that floating-point random values may be used as the rotation angles and translation amounts.

26 22 50 50 10 FIG. When zero is selected as all of the variation parameters, neither the posture nor the position is varied. In this case, in step Sdescribed later, the feature image generated in step Sis output without modification. That is, in the processing of, there is a case where the original feature image is output, without modification, as a feature image for training of the CNN, and there is also a case where an augmented image, in which at least one of the posture and the position is varied with respect to the original feature image, is generated and output as the feature image for training of the CNN.

25 124 23 24 [Step S] The data augmentation processing unitconverts the three-dimensional data calculated in step Ssuch that variation based on the variation parameters selected in step Sis applied to the vein pattern. 26 124 25 11 50 9 FIG. [Step S] The data augmentation processing unitconverts the three-dimensional data obtained in step Sinto two-dimensional data by using the camera matrix P in which the parameters calculated in step Sofare set. As a result, a feature image for training of the CNNis generated. Note that, in selection of the variation parameters, the probability of zero being selected may be set larger than the probabilities of the other discrete values being selected, such that the original feature image is more likely to be output as the feature image for training than the augmented image.

24 25 26 22 Note that, when zero has been selected as all the variation parameters in step S, the processing of step Sand the conversion processing using the camera matrix P in step Smay be skipped, and the feature image generated in step Smay be output as the feature image for training without modification.

10 FIG. 50 50 According to the processing ofdescribed above, variation of posture and position in a three-dimensional space is applied to the vein pattern in the feature image, and a feature image including the varied vein pattern is generated as an augmented image. As a result, it becomes possible to generate an augmented image that accurately represents three-dimensional variation of a palm that occurs in practice. Then, the CNNis trained using such augmented images as well, and the trained CNNis used in authentication processing, whereby biometric authentication that is robust to variation in posture and position becomes possible, and authentication accuracy is improved.

111 122 100 123 121 111 50 The learning images registered in the learning image databasemay be images that have already undergone normalization processing. In this case, the normalization processing unitof the learning apparatusis not needed, and the feature extracting unitexecutes the feature extraction processing on the learning images (normalized images) extracted by the training control unitfrom the learning image database. As a result, the processing load in training is reduced, and the training time of the CNNis shortened.

111 122 123 100 124 111 Further, the learning images registered in the learning image databasemay be feature images that have undergone normalization processing and feature extraction processing. In this case, the normalization processing unitand the feature extracting unitof the learning apparatusare not needed, and the data augmentation processing unitexecutes the data augmentation processing on the learning images (feature images) extracted from the learning image database.

200 50 Next, details of processing executed by the authentication apparatuswill be described. Thereby, the processing load in training is further reduced, and the training time of the CNNis shortened.

11 FIG. 200 210 221 222 223 224 225 illustrates an example of a configuration of processing functions included in the authentication apparatus. The authentication apparatusincludes a storing unit, a normalization processing unit, a feature extracting unit, an authentication data generating unit, a registration processing unit, and a matching processing unit.

210 200 113 100 210 211 210 211 The storing unitis a storage area secured in a storage device (not illustrated) included in the authentication apparatus. The weight datasetgenerated by the learning apparatusis stored in the storing unit. In addition, an authentication database (DB)is stored in the storing unit. In the authentication database, registration data corresponding to each user to be authenticated is registered in association with a user ID for identifying the user.

221 222 223 224 225 200 Processing of the normalization processing unit, the feature extracting unit, the authentication data generating unit, the registration processing unit, and the matching processing unitis implemented, for example, by a processor (not illustrated) included in the authentication apparatusexecuting a predetermined firmware program.

221 201 222 223 50 113 The normalization processing unitperforms normalization processing on a biometric image captured by the biometric sensor, and outputs a normalized image. The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the normalized image, and outputs a feature image. The authentication data generating unitinputs the feature image to the CNNconfigured using the weight dataset, and calculates a feature vector as authentication data.

224 211 225 211 225 211 225 The registration processing unit, at the time of registration processing, registers the calculated feature vector as registration data, together with a user ID, in the authentication database. The matching processing unit, at the time of authentication processing, compares the calculated feature vector serving as matching data with registration data in the authentication database, and outputs an authentication result. For example, when one-to-one authentication is performed, the matching processing registration data corresponding to a unitreads designated user ID from the authentication database, and calculates a matching score indicating similarity between the read registration data and the matching data. When the calculated matching score is equal to or greater than a predetermined threshold value, the matching processing unitoutputs an authentication result indicating successful authentication.

200 Next, processing of the authentication apparatuswill be described with reference to flowcharts.

12 FIG. 31 221 201 [Step S] The normalization processing unitacquires a captured image obtained by the biometric sensortogether with a user ID for identifying a user who is an imaging target. 32 221 21 10 FIG. [Step S] The normalization processing unitexecutes normalization processing on the captured image in the same manner as in step Sof, and generates a normalized image. 33 222 22 10 FIG. [Step S] The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the normalized image in the same manner as in step Sof, and generates a feature image. 34 223 50 113 [Step S] The authentication data generating unitinputs the feature image to the CNNin which each weight coefficient of the weight datasetis set, and calculates a feature vector. 35 224 211 31 [Step S] The registration processing unitregisters the calculated feature vector, as registration data, in the authentication databasetogether with the user ID acquired in step S. is a flowchart illustrating an example of registration processing.

13 FIG. 13 FIG. 41 221 201 [Step S] The normalization processing unitacquires a captured image obtained by the biometric sensortogether with a user ID. 42 221 21 10 FIG. [Step S] The normalization processing unitexecutes normalization processing on the captured image in the same manner as in step Sof, and generates a normalized image. 43 222 22 10 FIG. [Step S] The feature extracting unitexecutes feature extraction processing for extracting a vein pattern from the normalized image in the same manner as in step Sof, and generates a feature image. 44 223 50 113 [Step S] The authentication data generating unitinputs the feature image to the CNNin which each weight coefficient of the weight datasethas been set, and calculates a feature vector. 45 225 211 41 [Step S] The matching processing unitreads, from the authentication database, registration data corresponding to the user ID acquired in step S. 46 225 44 225 225 225 [Step S] The matching processing unitcalculates a matching score indicating similarity between the registration data that has been read out and the feature vector (matching data) calculated in step S. As the matching score, for example, a value inversely proportional to a distance between vectors of the registration data and the matching data, or a cosine similarity between the registration data and the matching data, is calculated. The matching processing unitcompares the calculated matching score with a predetermined threshold. When the matching score is equal to or greater than the threshold, the matching processing unitoutputs an authentication result indicating that authentication has succeeded, and when the matching score is less than the threshold, the matching processing unitoutputs an authentication result indicating that authentication has failed. is a flowchart illustrating an example of matching processing. In, as one example, it is assumed that one-to-one authentication is executed.

12 13 FIGS.and 50 100 201 In the processing ofdescribed above, authentication data (feature vectors) is calculated by the CNNthat has been trained by the learning apparatus, using both original learning images and augmented images. As a result, biometric authentication that is robust to variations in posture and position of a palm with respect to the biometric sensorbecomes possible, and authentication accuracy is improved.

24 10 FIG. Note that, in step Sof, as to the discrete values corresponding to each variation parameter, the same selection probability is set for the positive discrete values and the negative discrete values with respect to zero. However, for the translation amount in the z-axis direction among the variation parameters, these selection probabilities may be set in an asymmetric manner.

14 FIG. illustrates probabilities of setting discrete values for an amount of translation in the z-axis direction.

201 201 0 0 0 The value in the z-axis direction represents a distance from the biometric sensorto a palm, which is a biological body part. As described above, there exists Z, which indicates an appropriate distance from the biometric sensorto the palm. In a case where the palm is translated along the z-axis direction, when the distance to the palm is shorter than Z, a variation amount of the palm on an image becomes larger than when the distance is longer than Z. For example, even when the palm is rotated by the same angle in the same direction, the palm appears to be inclined to a greater extent on the image when the distance in the z-axis direction is smaller.

124 50 Accordingly, for the amount of translation in the z-axis direction, the data augmentation processing unitmay set a probability of selecting and setting (setting probability) of a discrete value greater than zero to be smaller than a setting probability of a discrete value less than zero, among settable discrete values. In this case, by probabilistically increasing the use of learning images in which apparent variation on the image is large and authentication is difficult, it becomes possible to generate learning images for the CNNthat enable improvement in authentication accuracy.

50 In front of a CNN for image classification, a spatial transformer network (STN) having a posture correction function may be disposed. In such a case, the STN and the CNN are capable of being trained simultaneously. Therefore, as a third embodiment, an example is described in which an STN is disposed in front of the CNNin the second embodiment.

15 FIG. 15 FIG. 200 223 223 60 50 a a a illustrates an example of a configuration of an authentication data generating unit included in an authentication apparatus according to the third embodiment. An authentication apparatusillustrated inincludes an authentication data generating unit. In the authentication data generating unit, an STNis disposed in front of the CNN.

in 223 60 50 60 221 60 50 a When an input image Iis input to the authentication data generating unit, the STNexecutes normalization processing for correcting posture to generate a normalized image, and the CNNgenerates authentication data (feature vector) based on the normalized image. That is, the STNrealizes the same function as the normalization processing unitin the second embodiment. In addition, the learning apparatus is capable of training the STNand the CNNsimultaneously.

16 FIG. 16 FIG. 7 FIG. 100 110 121 123 124 125 a a. illustrates an example of a configuration of processing functions of a learning apparatus according to the third embodiment. In, the same reference numerals are given to processing functions similar to those in. A learning apparatusaccording to the third embodiment includes the storing unit, the training control unit, the feature extracting unit, the data augmentation processing unit, and an STN-CNN training unit

110 111 112 113 113 60 50 a a The storing unitstores the learning image database, the sensor dataset, and a weight dataset. The weight datasetincludes weight coefficients included in the trained STNand CNN.

123 111 124 The feature extracting unitapplies feature extraction processing for extracting a vein pattern, not to normalized images but to learning images acquired from the learning image database, and outputs feature images to the data augmentation processing unit.

125 100 121 125 60 50 124 113 a a a a. Processing of the STN-CNN training unitis realized by a processor included in the learning apparatusexecuting a predetermined application program. Under control of the training control unit, the STN-CNN training unittrains both the STNand the CNNusing feature images for training obtained by the data augmentation processing unit, and generates the weight dataset

17 FIG. 17 FIG. 9 FIG. 12 14 15 17 12 14 15 17 a a a a 12 121 60 50 a [Step S] The training control unitinitializes weight coefficients of the STNand the CNN. For example, initial values of the weight coefficients are set by random values. 14 60 50 14 a 9 FIG. [Step S] Training data augmentation processing is executed. In this processing, each feature image for training of the STNand the CNNis generated on the basis of each of the predetermined number of learning images that have been extracted. Further, this processing differs from step Sofin that augmented images are generated from feature images in a state where normalization is not performed. 15 125 60 50 14 60 50 15 a a a 9 FIG. [Step S] The STN-CNN training unitexecutes training processing for the STNand the CNNby using each feature image generated in step S, and updates the weight coefficients of the STNand the CNN. In this processing, similarly to step Sof, each feature image is used as training data, and machine learning is performed in which ID of a person associated with the learning image serving as the source of the feature image is used as ground truth data. 17 125 110 113 60 50 a a a [Step S] The STN-CNN training unitstores, in the storing unitas the weight dataset, the weight coefficients that are set in the STNand the CNN. is a flowchart illustrating an example of training processing. In the training processing illustrated in, steps S, S, S, and Sare executed in place of steps S, S, S, and Sin the processing of, respectively.

18 FIG. 17 FIG. 18 FIG. 18 FIG. 10 FIG. 14 13 22 21 22 a a 22 123 22 a 10 FIG. [Step S] The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the learning image, and generates a feature image. The feature extraction processing may be executed in accordance with a procedure similar to that of step Sof. is a flowchart illustrating an example of training data augmentation processing. In step Sof, the processing inis executed for each of the learning images extracted in step S. In the training data augmentation processing illustrated in, step Sis executed in place of steps Sand Sof.

60 50 113 113 200 a a a Through the above processing, both the STNand the CNNare trained, and the weight datasetis generated as a result of training. The weight datasetis transferred to the authentication apparatusvia, for example, a network or a portable recording medium.

19 FIG. 19 FIG. 11 FIG. 200 210 222 223 224 225 a a illustrates an example of a configuration of processing functions included in an authentication apparatus according to the third embodiment. In, the same reference numerals are used to denote the same processing functions as those illustrated in. The authentication apparatusaccording to the third embodiment includes the storing unit, the feature extracting unit, an authentication data generating unit, the registration processing unit, and the matching processing unit.

211 113 100 210 222 201 223 a a a. The authentication databaseand the weight datasetgenerated by the learning apparatusare stored in the storing unit. The feature extracting unitperforms feature extraction processing for extracting a vein pattern not from a normalized image but from a biological image captured by the biometric sensor, and outputs a feature image to the authentication data generating unit

223 200 223 60 113 223 50 113 a a a a a a Processing of the authentication data generating unitis implemented, for example, by a processor included in the authentication apparatusexecuting a predetermined firmware program. The authentication data generating unitinputs the feature image to the STNformed using the weight dataset, and normalizes the feature image. Further, the authentication data generating unitinputs the normalized image to the CNNformed using the weight dataset, and calculates a feature vector as authentication data.

20 FIG. 20 FIG. 12 FIG. 33 34 34 2 32 34 a al a 33 222 31 a [Step S] The feature extracting unitperforms feature extraction processing for extracting a vein pattern from the captured image acquired in step S, and generates a feature image. 34 1 223 60 60 113 a a a [Step S] The authentication data generating unitinputs the feature image to the STNin which weight coefficients for the STNincluded in the weight datasetare set, normalizes the feature image, and generates a normalized image. 34 2 223 50 50 113 a a a [Step S] The authentication data generating unitinputs the normalized image to the CNNin which weight coefficients for the CNNincluded in the weight datasetare set, and calculates a feature vector. is a flowchart illustrating an example of registration processing. In the registration processing illustrated in, steps S, S, and Sare executed in place of steps Sto Sin the processing of.

224 211 The calculated feature vector is input to the registration processing unitand is registered, as registration data, in the authentication databasetogether with the user ID.

200 31 34 2 45 a a 20 FIG. 13 FIG. Although the illustration is omitted, in matching processing performed by the authentication apparatus, a feature vector as matching data is calculated by processing similar to the processing from step Sto step Sof. Then, the matching data is matched with the registration data by the processing of step Sof, and an authentication result is output.

50 60 60 60 According to the third embodiment described above, not only the CNNbut also the STNare trained using learning images and augmented images. Accordingly, the normalization processing by the STNis also optimized by augmented images obtained by applying three-dimensional variation to the learning images, thereby making it possible to generate the STNthat is capable of performing accurate normalization processing with respect to posture variations.

21 FIG. 21 FIG. 100 300 410 420 300 410 420 b illustrates an example of a configuration of a biometric authentication system according to a fourth embodiment. The biometric authentication system illustrated inincludes a learning apparatus, an authentication server, and authentication terminalsand. Among these components, at least the authentication serverand the authentication terminalsandare connected via a network.

200 300 410 420 410 401 401 300 300 420 402 402 300 300 In this biometric authentication system, the processing functions of the authentication apparatusdescribed above are implemented by the authentication serverand the authentication terminalsand, which are client devices, in a client-server manner. The authentication terminalincludes a biometric sensor, and a biometric image captured by the biometric sensoris transmitted to the authentication server, and data registration processing and data matching processing are executed in the authentication server. Similarly, the authentication terminalincludes a biometric sensor, and a biometric image captured by the biometric sensoris transmitted to the authentication server, and data registration processing and data matching processing are executed in the authentication server.

300 50 200 100 300 b The authentication servergenerates authentication data (feature vectors) by using the CNN, similarly to the authentication apparatusdescribed above. The learning apparatusperforms training processing of the CNN and transfers an obtained weight dataset to the authentication server.

401 410 402 420 401 402 100 50 401 402 b Here, it is assumed that the specification of the biometric sensorof the authentication terminaldiffers from the specification of the biometric sensorof the authentication terminal. In such a case, even when the same biological body part is captured in the same posture and at the same distance by both biometric sensorsand, the appearance of the biological body part in the captured images differs. Therefore, when the learning apparatusgenerates augmented images for training of the CNN, it becomes appropriate to use camera matrices P that correspond to the sensor specifications for the biometric sensorand the biometric sensor.

100 401 402 100 401 50 401 100 402 50 402 b b b Accordingly, the learning apparatusholds sensor datasets corresponding to the biometric sensorsand, respectively. The learning apparatusgenerates augmented images by using the sensor dataset for the biometric sensor, and generates the CNNfor the biometric sensorby using the original learning images and the augmented images. In addition, the learning apparatusgenerates augmented images by using the sensor dataset for the biometric sensor, and generates the CNNfor the biometric sensorby using the original learning images and the augmented images.

50 401 50 402 100 300 300 401 402 b Then, the weight dataset of the CNNfor the biometric sensorand the weight dataset of the CNNfor the biometric sensorare transferred from the learning apparatusto the authentication server. The authentication serverstores each weight dataset in association with a sensor ID for identifying the sensor type of the biometric sensorsand.

300 410 401 300 50 50 300 420 402 300 50 50 The authentication serverreceives, from the authentication terminal, a captured image together with a sensor ID corresponding to the biometric sensor. The authentication serverforms the CNNby using weight parameters associated with the received sensor ID, and generates authentication data by the CNN. Similarly, the authentication serverreceives, from the authentication terminal, a captured image together with a sensor ID corresponding to the biometric sensor. The authentication serverforms the CNNby using weight parameters associated with the received sensor ID, and generates authentication data by the CNN.

401 402 100 50 b Through such processing, highly accurate biometric authentication becomes possible regardless of differences in specifications, by using biometric images captured by biometric sensorsandhaving different specifications. For example, even when the same three-dimensional variation is applied to a biological body part, if the specifications of the biometric sensor that captures the biometric body part are different, the appearance of the biological body part after variation in the captured images also differs. As one example, when the focal lengths of the biometric sensors differ, the appearance such as the size of the biological body part in the captured images also differs. According to the training processing by the learning apparatusdescribed above, such differences in appearance according to differences in specifications of biometric sensors become reproducible through training. Therefore, by calculating authentication data using the CNNtrained for each biometric sensor specification, highly accurate authentication processing that is robust to three-dimensional variations of biological body parts is achieved.

22 FIG. 22 FIG. 7 FIG. 100 110 121 122 123 124 125 121 100 b b b b illustrates an example of a configuration of processing functions included in the learning apparatus. In, the same reference numerals are assigned to processing functions similar to those in. The learning apparatusincludes the storing unit, a training control unit, the normalization processing unit, the feature extracting unit, the data augmentation processing unit, and the CNN training unit. The processing of the training control unitis implemented, for example, by a processor included in the learning apparatusexecuting a predetermined application program.

110 111 112 1 112 2 113 1 113 2 112 1 401 401 112 2 402 402 113 1 401 113 2 402 b b b b b b b b The storing unitstores the learning image database, sensor datasetsand, and weight datasetsand. The sensor datasetis associated with a sensor ID of the biometric sensor, and camera parameters corresponding to the sensor type of the biometric sensorare registered in advance. The sensor datasetis associated with a sensor ID of the biometric sensor, and camera parameters corresponding to the sensor type of the biometric sensorare registered in advance. The weight datasetis associated with the sensor ID of the biometric sensor, and the weight datasetis associated with the sensor ID of the biometric sensor.

100 121 122 123 124 125 121 124 124 125 110 b b b In the learning apparatus, under control of the training control unit, the normalization processing unit, the feature extracting unit, the data augmentation processing unit, and the CNN training unitgenerate weight datasets to be set in the CNN basically in the same procedure as in the second embodiment. However, the training control unitdesignates a sensor ID to the data augmentation processing unit, and the data augmentation processing unitgenerates the camera matrix P based on a sensor dataset corresponding to the designated sensor ID, and generates augmented images by using the camera matrix P. When a weight dataset of the CNN is generated by the training processing, the CNN training unitstores the weight dataset in the storing unitwith the sensor ID associated with the weight dataset.

23 FIG. illustrates an example of a configuration of processing functions included in the authentication server and the authentication terminals.

410 430 440 430 410 430 431 401 440 410 440 401 431 300 440 300 a a a a a a a a a First, the authentication terminalincludes a storing unitand a communicating unit. The storing unitis a storage area secured in a storage device included in the authentication terminal. The storing unitstores sensor ID datain which a sensor ID of the biometric sensoris described. The processing of the communicating unitis implemented, for example, by a processor included in the authentication terminalexecuting a predetermined firmware program. The communicating unittransmits, at the time of data registration and at the time of authentication, a biometric image captured by the biometric sensor, a user ID, and the sensor ID described in the sensor ID data, to the authentication server. In addition, the communicating unitreceives an authentication result from the authentication serverat the time of authentication, and outputs the authentication result via a display device or the like.

420 430 440 430 431 402 440 402 431 300 440 300 b b b b b b b Similarly, the authentication terminalincludes a storing unitand a communicating unit. The storing unitstores sensor ID datain which a sensor ID of the biometric sensoris described. The communicating unittransmits, at the time of data registration and at the time of authentication, a biometric image captured by the biometric sensor, a user ID, and the sensor ID described in the sensor ID data, to the authentication server. In addition, the communicating unitreceives an authentication result from the authentication serverat the time of authentication, and outputs the authentication result via a display device or the like.

300 310 311 312 313 314 315 316 The authentication serverincludes a storing unit, a communicating unit, a normalization processing unit, a feature extracting unit, an authentication data generating unit, a registration processing unit, and a matching processing unit.

310 300 310 113 1 113 2 100 311 311 311 401 311 402 b b b a b a b The storing unitis a storage area secured in a storage device included in the authentication server. The storing unitstores weight datasetsandgenerated by the learning apparatus, and authentication databases (DBs)and. The authentication databaseis associated with a sensor ID of the biometric sensor, and the authentication databaseis associated with a sensor ID of the biometric sensor.

311 312 313 314 315 316 300 The processing of the communicating unit, the normalization processing unit, the feature extracting unit, the authentication data generating unit, the registration processing unit, and the matching processing unitis implemented, for example, by a processor included in the authentication serverexecuting a predetermined application program.

311 410 420 311 410 420 The communicating unitreceives, at the time of data registration and at the time of authentication, biometric images, user IDs, and sensor IDs from the authentication terminalsand. In addition, the communicating unittransmits authentication results to the authentication terminalsandat the time of authentication.

312 313 314 315 316 221 222 223 224 225 314 50 311 50 315 316 311 11 FIG. The normalization processing unit, the feature extracting unit, the authentication data generating unit, the registration processing unit, and the matching processing unitbasically execute processing similar to that of the normalization processing unit, the extracting feature unit, the authentication data generating unit, the registration processing unit, and the matching processing unitin, respectively. However, the authentication data generating unitforms the CNNusing a weight dataset corresponding to the sensor ID received by the communicating unit, and calculates authentication data (feature vectors) by using the CNN. The registration processing unitregisters registration data in an authentication database corresponding to the sensor ID. The matching processing unitextracts registration data from an authentication database corresponding to the sensor ID, matches the extracted registration data with matching data, and transmits an authentication result to the communicating unit.

312 410 420 410 420 300 312 313 410 420 410 420 300 The normalization processing unitmay be mounted in the authentication terminalsand. In this images are transmitted from the case, normalized authentication terminalsandto the authentication servertogether with user IDs and sensor IDs. In addition, the normalization processing unitand the feature extracting unitmay be mounted in the authentication terminalsand. In this case, feature images are transmitted from the authentication terminalsandto the authentication servertogether with user IDs and sensor IDs.

312 313 314 410 420 113 1 430 410 314 410 50 113 1 440 410 300 b a b a Further, the normalization processing unit, the feature extracting unit, and the authentication data generating unitmay be mounted in the authentication terminalsand. In this case, the weight datasetis stored in the storing unitof the authentication terminal, and the authentication data generating unitof the authentication terminalcalculates authentication data (feature vectors) by using the CNNformed by using the weight dataset. The communicating unitof the authentication terminaltransmits the calculated authentication data to the authentication servertogether with a user ID and a sensor ID.

113 2 430 420 314 420 50 113 2 440 420 300 b b b b In addition, the weight datasetis stored in the storing unitof the authentication terminal, and the authentication data generating unitof the authentication terminalcalculates authentication data (feature vectors) by using the CNNformed by using the weight dataset. The communicating unitof the authentication terminaltransmits the calculated authentication data to the authentication servertogether with a user ID and a sensor ID.

10 100 100 100 200 200 300 410 420 a b a The processing functions of the apparatuses described in the embodiments above (for example, the learning apparatuses,,, and, the authentication apparatusesand, the authentication server, and the authentication terminalsand) are implementable by a computer. In this case, a program describing processing contents of the functions to be included in each apparatus is provided, and by executing the program by a computer, the processing functions described above are implemented on the computer. The program describing the processing contents is recordable on a computer-readable recording medium. Examples of the computer-readable recording medium include magnetic storage devices, optical discs, and semiconductor memories. Examples of the magnetic storage devices include hard disk drives (HDDs) and magnetic tapes. Examples of the optical discs include compact discs (CDs), digital versatile discs (DVDs), and Blu-ray discs (BDs, registered trademark).

When the program is distributed, for example, a portable recording medium such as a DVD or a CD on which the program is recorded is sold. In addition, the program may be stored in a storage device of a server computer and transferred from the server computer to another computer via a network.

A computer that executes the program stores, in its own storage device, the program recorded on the portable recording medium or the program transferred from the server computer. Then, the computer reads the program from its own storage device and executes processing according to the program. The computer may also directly read the program from the portable recording medium and execute processing according to the program. In addition, the computer may sequentially execute processing according to a received program each time the program is transferred from a server computer connected via a network.

In one aspect, a neural network that enables biometric authentication robust to variations in the posture of a biological body part is generated.

All examples and conditional language provided herein are intended for the pedagogical purposes of aiding the reader in understanding the invention and the concepts contributed by the inventor to further the art, and are not to be construed as limitations to such specifically recited examples and conditions, nor does the organization of such examples in the specification relate to a showing of the superiority and inferiority of the invention. Although one or more embodiments of the present invention have been described in detail, it should be understood that various changes, substitutions, and alterations could be made hereto without departing from the spirit and scope of the invention.

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

March 31, 2026

Publication Date

August 6, 2026

Inventors

Takahiro AOKI

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “LEARNING METHOD AND LEARNING APPARATUS” (US-20260229023-A1). https://patentable.app/patents/US-20260229023-A1

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.