Patentable/Patents/US-12675673-B2
US-12675673-B2

Multidimensional data processing device, method, and computer-readable recording medium

PublishedJuly 7, 2026
Assigneenot available in USPTO data we have
InventorsSeiya Shibata
Technical Abstract

92 93 95 96 95 98 The 1 dimensional data generation meansgenerates 1 dimensional data, by setting the number of elements of each dimension other than predetermined one dimension to 1 based on multidimensional data corresponding to one input data. The number of elements reducing meansreduces the number of elements included in the 1 dimensional data. The copy meansgenerates multidimensional data, by copying the 1 dimensional data whose number of elements has been reduced multiple times. The convolution layer processing meansperforms a convolution layer process with a filter size of 1×1 on the multidimensional data generated by the copy means. The element-wise product operation meansperforms an element-wise product operation, based on the multidimensional data corresponding to one input data and multidimensional data generated by the above process.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

a memory configured to store instructions; and generate by setting a number of elements of each dimension other than a predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1-dimensional data whose number of elements is the same as a number of elements of the predetermined one dimension in the multidimensional data; reduce the number of elements included in the 1-dimensional data; copy, multiple times, the 1-dimensional data whose number of elements has been reduced, thereby to generate second multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; perform a convolution layer process with a filter size of 1×1 on the second multidimensional data, thereby to generate third multidimensional data in which the number of elements of the predetermined one dimension in the second multidimensional data is restored to original number of elements; and at least one processor configured to execute the instructions to: wherein a number of elements of the 1-dimensional data whose number of elements has been reduced and the number of elements of the predetermined one dimension in the second multidimensional data are the same, and wherein when A denotes the number of elements of the 1-dimensional data whose number of elements has been reduced, i denotes an integer from 0 to A-1, each i-th element of the predetermined one dimension in the second multidimensional data is the same as i-th element of the predetermined one dimension in the 1-dimensional data whose number of elements has been reduced. perform an element-wise product operation, based on the multidimensional data corresponding to one input data and the third multidimensional data, . A multidimensional data processing device comprising:

2

claim 1 wherein the at least one processor is configured to apply a sigmoid function to each element in the third multidimensional data; wherein the at least one processor is configured to perform the element-wise product operation of the multidimensional data corresponding to one input data and the third multidimensional data after the sigmoid function is applied to each element. . The multidimensional data processing device according to,

3

claim 1 wherein the at least one processor is configured to change values of the elements in the 1-dimensional data whose number of elements has been reduced that have negative values to 0; wherein the at least one processor is configured to copy the 1-dimensional data multiple times. . The multidimensional data processing device according to,

4

claim 1 wherein the multidimensional data corresponding to one input data is 3 dimensional data, and, the predetermined one dimension is a dimension of channel. . The multidimensional data processing device according to,

5

generating, by setting a number of elements of each dimension other than a predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1-dimensional data whose number of elements is the same as a number of elements of the predetermined one dimension in the multidimensional data; reducing the number of elements included in the 1-dimensional data; executing a copy process of copying, multiple times, the 1-dimensional data whose number of elements has been reduced, thereby to generate second multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; performing a convolution layer process with a filter size of 1×1 on the second multidimensional data generated by the copy process, thereby to generate third multidimensional data in which the number of elements of the predetermined one dimension in the second multidimensional data generated by the copy process is restored to original number of elements; and performing an element-wise product operation, based on the multidimensional data corresponding to one input data and the third multidimensional data whose number of elements in the predetermined one dimension is restored to original number of elements, wherein a number of elements of the 1-dimensional data whose number of elements has been reduced and the number of elements of the predetermined one dimension in the second multidimensional data are the same, and wherein when A denotes the number of elements of the 1-dimensional data whose number of elements has been reduced, i denotes an integer from 0 to A-1, each i-th element of the predetermined one dimension in the second multidimensional data is the same as i-th element of the predetermined one dimension in the 1-dimensional data whose number of elements has been reduced. . A multidimensional data processing method comprising:

6

claim 5 applying a sigmoid function to each element in the third multidimensional data whose number of elements of the predetermined one dimension is restored to original number of elements, wherein performing the element-wise product operation includes performing the element-wise product operation of the multidimensional data corresponding to one input data and the third multidimensional data after the sigmoid function is applied to each element. . The multidimensional data processing method according to, further comprising:

7

a 1-dimensional data generation process of generating, by setting a number of elements of each dimension other than a predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1-dimensional data whose number of elements is the same as a number of elements of the predetermined one dimension in the multidimensional data; a number of elements reducing process of reducing the number of elements included in the 1-dimensional data; a copy process of copying, multiple times, the 1-dimensional data whose number of elements has been reduced, thereby to generate second multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; a multidimensional data generation process of performing a convolution layer process with a filter size of 1×1 on the second multidimensional data generated by the copy process, thereby to generate third multidimensional data in which the number of elements of the predetermined one dimension in the second multidimensional data generated by the copy process is restored to original number of elements; and an element-wise product operation process of performing an element-wise product operation, based on the multidimensional data corresponding to one input data and the third multidimensional data generated by the multidimensional data generation process, wherein a number of elements of the 1-dimensional data whose number of elements has been reduced and the number of elements of the predetermined one dimension in the second multidimensional data are the same, and wherein when A denotes the number of elements of the 1-dimensional data whose number of elements has been reduced, i denotes an integer from 0 to A-1, each i-th element of the predetermined one dimension in the second multidimensional data is the same as i-th element of the predetermined one dimension in the 1-dimensional data whose number of elements has been reduced. . A non-transitory computer-readable recording medium in which a multidimensional data processing program is recorded, wherein the multidimensional data processing program causes a computer to execute:

8

claim 7 a sigmoid function process of applying a sigmoid function to each element in the third multidimensional data generated by the multidimensional data generation process, wherein the multidimensional data processing program causes the computer to execute in the element-wise product operation process, performing the element-wise product operation of the multidimensional data corresponding to one input data and the third multidimensional data after the sigmoid function is applied to each element. wherein the multidimensional data processing program causes the computer to execute: . The non-transitory computer-readable recording medium in which the multidimensional data processing program is recorded according to,

Detailed Description

Complete technical specification and implementation details from the patent document.

This application is a National Stage Entry of PCT/JP2020/020567 filed on May 25, 2020, the contents of all of which are incorporated herein by reference, in their entirety.

The present invention relates to a neural network processing device and a neural network processing method that perform a process of a neural network, and a computer-readable recording medium recording a neural network processing program.

In a neural network, a block is a batch of multiple layers which are basic components.

1 1 1 8 FIG. 9 FIG. NPLdescribes a SE (Squeeze-and-Excitation) block, as a block that improves the accuracy of CNN (Convolutional Neural Network).is a schematic diagram showing the SE block described in NPL. NPLshows the case where 3 dimensional data U is input to the SE block.is a schematic diagram showing the 3 dimensional data U input to the SE block.

The individual dimensions in the 3 dimensional data are referred to as the H dimension, the W dimension, and the C dimension. The H dimension is, for example, the dimension related to the height of an image. The W dimension is, for example, the dimension related to the width of the image. The C dimension is the dimension related to the channel. It is assumed that the number of elements of the H dimension in the 3 dimensional data U is H. It is assumed that the number of elements of the W dimension in the 3 dimensional data U is W. It is assumed that the number of elements of the C dimension in the 3 dimensional data U is C. The size of the 3 dimensional data U can be expressed as H×W×C.

101 10 FIG. In the Global Pooling layer (step S), the number of elements of the H dimension and the W dimension are respectively 1. The number of elements of the C dimension remains unchanged at C. In other words, based on the 3 dimensional data U whose size is H×W×C, 1 dimensional data whose size is 1×1×C is generated.is a schematic diagram showing the 1 dimensional data obtained in the Global Pooling layer.

102 11 FIG. In the first FC (Fully Connected) layer (step S), the number of elements in the 1 dimensional data obtained in the Global Pooling layer is reduced.is a schematic diagram showing the 1 dimensional data obtained in the FC layer. Here, the number of elements after the reduction is A. A<C.

12 FIG. 12 FIG. 11 FIG. 102 102 is a schematic diagram showing the process in the first FC layer (step S). In the first FC layer (step S), the number of elements that are outputs is less than the number of elements that are inputs. Then, the elements that are inputs and the elements that are outputs are fully connected as shown in, and weights are determined for respective individual connections. When the number of elements that are inputs is C and the number of elements that are outputs is A, the number of weights is C×A. Each weight is determined in advance by learning. The value of an element that is an output is calculated based on the values of the individual elements that are inputs connected with the element and the weights determined for each pair of the element that is an output and the individual element that is an input. By finding the values of the A elements that are outputs, 1 dimensional data (see) whose number of elements is A is obtained.

103 102 In the ReLU (Rectified Linear Unit) layer (Step S), among the elements in the 1 dimensional data obtained in the FC layer (Step S), the values of elements with negative values are changed to 0. The values of elements with values equal to or greater than 0 are not changed. In the ReLU layer, the number of elements in 1 dimensional data remains unchanged at A.

104 In the second FC layer (step S), the number of elements in the 1 dimensional data obtained in the ReLU layer is increased back to the original number of elements (C elements).

13 FIG. 13 FIG. 104 104 is a schematic diagram showing the process in the second FC layer (step S). In the second FC layer (step S), the number of elements that are outputs is greater than the number of elements that are inputs. Then, the elements that are inputs and the elements that are outputs are fully connected as shown in, and weights are determined for respective individual connections. When the number of elements that are inputs is A and the number of elements that are outputs is C, the number of weights is A×C. Each weight is determined in advance by learning. The value of an element that is an output is calculated based on the values of the individual elements that are inputs connected with the element and the weights determined for each pair of the element that is an output and the individual element that is an input. By finding the values of the C elements that are outputs, 1 dimensional data whose number of elements is C is obtained.

The first FC layer and the second FC layer differ only in whether the number of elements that are outputs decreases or increases with respect to the number of elements that are inputs; the essential process is the same.

105 In the Sigmoid layer (step S), the sigmoid function is applied to each element in the 1 dimensional data obtained in the second FC layer. In the Sigmoid layer, the number of elements in the 1 dimensional data remains unchanged at C.

Individual elements in the 1 dimensional data obtained by the Sigmoid layer are used as coefficients representing the degree of importance of the channel corresponding to the individual element. For example, the 0th element in the 1 dimensional data is the coefficient representing the degree of importance of the 0th channel.

106 9 FIG. 14 FIG. 15 FIG. In the Scale layer (step S), the elements of each channel of the first input 3 dimensional data U (see) are multiplied by a coefficient indicating the degree of importance of that channel. At this time, by copying the 1 dimensional data obtained in the Sigmoid layer H×W times, 3 dimensional data whose size is H×W×C is generated. This 3 dimensional data is denoted by a symbol X′. Since the size of the 1 dimensional data obtained in the Sigmoid layer is 1×1×C, by copying this 1 dimensional data H×W times, 3 dimensional data X′ with size H×W×C is obtained.is a schematic diagram showing the 3 dimensional data X′ obtained by copying the 1 dimensional data of size 1×1×C, H×W times.is a schematic diagram showing calculation of element-wise product of the 3 dimensional data U and the 3 dimensional data X′. The sizes of both the 3 dimensional data U and the 3 dimensional data X′ are H×W×C and are common. Furthermore, the elements in the 3 dimensional data U and the elements in the 3 dimensional data X′ can both be specified by 3 dimension coordinates. Therefore, it is possible to associate elements in the 3 dimensional data U and elements in the 3 dimensional data X′ that share the same 3 dimension coordinates. As a result, the elements in the 3 dimensional data U and the elements in the 3 dimensional data X′ are associated one-to-one. By calculating the product of the values of elements for each pair of elements to be associated, new 3 dimensional data whose size is H×W×C is obtained. This 3 dimensional data is the result of the element-wise product of the 3 dimensional data U and the 3 dimensional data X′, and is the output of the Scale layer. The 3 dimensional data obtained by this element-wise product operation can be said to be the data obtained by multiplying the multiple elements for each individual channel of the 3 dimensional data U by the coefficient corresponding to the channel (coefficient representing the degree of importance).

The output of the Scale layer (element-wise product of the 3 dimensional data U and the 3 dimensional data X′) is also the output of the SE block.

NPL 1: Jie Hu, Li Shen, Samuel Albanie, Gang Sun, Enhua Wu, “Squeeze-and-Excitation Networks”, [online], [retrieved Apr. 3, 2020], Internet<URL: https://arxiv.org/pdf/1709.01507.pdf>

The SE block can improve the accuracy of CNN. However, the SE block can significantly reduce processing speed.

Therefore, it is the object of the present invention to provide a neural network processing device, a neural network processing method, and a computer-readable recording medium recording a neural network processing program that can obtain the same CNN accuracy as the SE block and perform operations faster than the SE block.

A neural network processing device according to the present invention includes: 1 dimensional data generation means for generating, by setting the number of elements of each dimension other than predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1 dimensional data whose number of elements is the same as the number of elements of the predetermined one dimension in the multidimensional data; number of elements reducing means for reducing the number of elements included in the 1 dimensional data; copy means for copying the 1 dimensional data whose number of elements has been reduced multiple times, thereby to generate multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; convolution layer processing means for performing a convolution layer process with a filter size of 1×1 on the multidimensional data generated by the copy means, thereby to generate multidimensional data in which the number of element of the predetermined one dimension in multidimensional data generated by the copy means is restored to original number of elements; and element-wise product operation means for performing an element-wise product operation, based on the multidimensional data corresponding to one input data and the multidimensional data generated by the convolution layer processing means.

A neural network processing method according to the present invention includes: generating, by setting the number of elements of each dimension other than predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1 dimensional data whose number of elements is the same as the number of elements of the predetermined one dimension in the multidimensional data; reducing the number of elements included in the 1 dimensional data; executing a copy process of copying the 1 dimensional data whose number of elements has been reduced multiple times, thereby to generate multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; performing a convolution layer process with a filter size of 1×1 on the multidimensional data generated by the copy process, thereby to generate multidimensional data in which the number of element of the predetermined one dimension in multidimensional data generated by the copy process is restored to original number of elements;

and performing an element-wise product operation, based on the multidimensional data corresponding to one input data and the multidimensional data whose number of element of the predetermined one dimension is restored to original number of elements.

A computer-readable recording medium according to the present invention is a computer-readable recording medium in which a neural network processing program is recorded, wherein the neural network processing program causes a computer to execute: a 1 dimensional data generation process of generating, by setting the number of elements of each dimension other than predetermined one dimension to 1 based on multidimensional data corresponding to one input data, 1 dimensional data whose number of elements is the same as the number of elements of the predetermined one dimension in the multidimensional data; a number of elements reducing process of reducing the number of elements included in the 1 dimensional data; a copy process of copying the 1 dimensional data whose number of elements has been reduced multiple times, thereby to generate multidimensional data in which the number of elements of each dimension other than the predetermined one dimension is restored to original number of elements; a multidimensional data generation process of performing a convolution layer process with a filter size of 1×1 on the multidimensional data generated by the copy process, thereby to generate multidimensional data in which the number of element of the predetermined one dimension in multidimensional data generated by the copy process is restored to original number of elements; and an element-wise product operation process of performing an element-wise product operation, based on the multidimensional data corresponding to one input data and the multidimensional data generated by the multidimensional data generation process.

According to this invention, it is possible to obtain the same CNN accuracy as the SE block and perform operations faster than the SE block.

As mentioned above, the SE block can improve the accuracy of CNN. However, the SE block can significantly reduce processing speed.

The inventor of the present invention considered the following reasons for the reduced processing speed when SE block is used.

14 FIG. As mentioned above, in the SE block, in the Scale layer, the 1 dimensional data whose size is 1×1×C is copied H×W times to obtain the 3 dimensional data X′ (see) whose size is H×W×C. This H×W times copy process causes a large overhead.

In particular, when the number of elements in the 1 dimensional data is large (in other words, when the value of C is large), the number of times an element is read from a memory and written to memory becomes enormous, and therefore the overhead of H×W times copy process is also enormous. For example, it is assumed that the size of the 1 dimensional data obtained in the Sigmoid layer is 1×1×1024 (i.e. C=1024). It is assumed that the size of the 3 dimensional data U is 7×7×1024. In other words, H=7 and W=7. In this case, for each of the 1024 elements in the 1 dimensional data, the read and write processes must be performed 7×7=49 times, resulting in a very large overhead due to the copy process.

The inventor of the present invention considered that the large overhead caused by this copy process was the cause of the slow processing speed in the SE block. Based on this consideration, the inventor made the following invention.

Example embodiment of the present invention is described below with reference to the drawings.

Multidimensional data is input to the neural network processing device of the example embodiment of the present invention. One multidimensional data corresponds to one input data. In order to make the invention easier to understand, the case in which the multidimensional data corresponding to one input data is 3 dimensional data will be used as an example in the present example embodiment. Even if the multidimensional data is other than 3 dimensional data, the same processing as in the following example embodiment can be applied.

9 FIG. The multidimensional data (3 dimensional data in the present example embodiment) input to the neural network processing device in the present example embodiment is denoted by the symbol U. As in the previous case, the individual dimensions in the 3 dimensional data U will be referred to as the H dimension, the W dimension, and the C dimension. The H dimension is, for example, the dimension related to the height of an image. The W dimension is, for example, the dimension related to the width of the image. The C dimension is the dimension related to the channel. It is assumed that the number of elements of the H dimension in the 3 dimensional data U is H. It is assumed that the number of elements of the W dimension in the 3 dimensional data U is W. It is assumed that the number of elements of the C dimension in the 3 dimensional data U is C. The size of the 3 dimensional data U can be expressed as H×W×C. The 3 dimensional data U can be represented schematically as shown in.

1 FIG. 1 2 3 4 5 6 7 8 is a block diagram showing an example configuration of a neural network processing device of the example embodiment of the present invention. A neural network processing deviceof the present example embodiment includes a 1 dimensional data generation unit, a FC layer processing unit, a ReLU layer processing unit, a copy unit, a convolution layer processing unit, a Sigmoid layer processing unit, and a Scale layer processing unit.

2 2 The 3 dimensional data U corresponding to one input data is input to the 1 dimensional data generation unit. Then, the 1 dimensional data generation unit, by setting the number of elements of each dimension other than the predetermined one dimension among the three dimensions (H dimension, W dimension, and C dimension) to 1, generates 1 dimensional data whose number of elements is the same as the number of elements of the predetermined one dimension in the 3 dimensional data U.

2 2 10 FIG. In the present example embodiment, it is assumed that the above predetermined one dimension is the “C dimension”. In this case, the 1 dimensional data generation unitsets the number of elements of each dimension other than the C dimension (H dimension and W dimension) to 1, thereby generates the 1 dimensional data whose number of elements is the same as the number of elements of the C dimension (i.e., C elements) in the 3 dimensional data U. The size of this 1 dimensional data is 1×1×C. The 1 dimensional data generated by the 1 dimensional data generation unitcan be represented schematically as shown in.

2 2 2 2 2 2 The 1 dimensional data generation unitmay, for example, generate 1 dimensional data by performing the same process as the Global Pooling layer. For example, the 1 dimensional data generation unitmay obtain the average value (which may be the maximum value) of the H×w elements in the 0th channel of the 3 dimensional data U, and determine the value as the value of the 0th element in the 1 dimensional data. The 1 dimensional data generation unitperforms the same process for each channel after the first, and determines the value of the element corresponding to each channel. As a result, C elements from the 0th to the C-1st are obtained, and 1 dimensional data with those elements is obtained. In this process, the 1 dimensional data generation unitsets the number of elements of the H dimension and the number of elements of the W dimension to 1. The size of the 1 dimensional data obtained in this process is 1×1×C. The number of elements of the H dimension and the number of elements of the W dimension are both 1. Here, a case in which the 1 dimensional data generation unitgenerates 1 dimensional data by performing the same process as it of the Global Pooling layer is described, but the 1 dimensional data generation unitmay generate 1 dimensional data whose size is 1×1×C by other methods.

3 2 The FC layer processing unitreduces the number of elements in 1 dimensional data by performing FC layer processing on the 1 dimensional data obtained by 1 dimensional data generation unit. As in the previous case, the number of elements after the reduction is A. A<C.

3 102 3 2 3 3 3 3 8 FIG. 12 FIG. 11 FIG. The processing of FC layer processing unitis the same as the processing of the first FC layer of the SE block (step Sin). That is, the FC layer processing unittakes the C elements in the 1 dimensional data obtained by 1 dimensional data generation unitas the elements that are inputs. In addition, the FC layer processing unittakes the A elements after the number of elements is reduced as the elements that are outputs. The elements that are inputs and the elements that are outputs are fully connected as shown in, and the weights of each connection are determined in advance by learning. In this case, C×A weights have been determined by learning in advance. The FC layer processing unitcalculates one element in A elements based on the values of each element of C and the weights determined for each pair of the one element and each of C elements. The FC layer processing unitcalculates values of each of A elements, to derive the 1 dimensional data whose number of elements is A. The size of this 1 dimensional data is 1×1×A. The 1 dimensional data derived by the FC layer processing unitcan be schematically represented as shown in.

3 It could be said that the FC layer processing unitperforms a process to reduce the number of elements in the 1 dimensional data.

4 3 4 4 4 1 4 4 1 The ReLU layer processing unitchanges the values of the elements in the 1 dimensional data derived by the FC layer processing unitthat have negative values to 0. The ReLU layer processing unitdoes not change the values of the elements whose values are equal to or greater than 0. The number of elements in the 1 dimensional data remains unchanged at A by the process of the ReLU layer processing unit. The ReLU layer processing unitmay not be included in the neural network processing device, and the above processing by the ReLU layer processing unitmay be omitted. Also, instead of the ReLU layer processing unit, a component that applies an arbitrary activation function to 1 dimensional data may be included in the neural network processing device.

5 4 5 The copy unitcopies the 1 dimensional data after processing by ReLU layer processing unitmultiple times (more specifically, H×W times) to generate multidimensional data in which the number of elements of each dimension other than the predetermined one dimension (C dimension) is restored to the original number of elements. Here, the “original number of elements” is the number of elements in the input multidimensional data (3 dimensional data U in the present example embodiment). In other words, by copying 1 dimensional data whose size is 1×1×A, H×W times, the copy unitgenerates 3 dimensional data in which the number of elements of the H dimension is H, the number of elements of the W dimension is W, and the number of elements of the C dimension is A. The size of this 3 dimensional data is H×W×A.

2 FIG. Hereafter, this 3 dimensional data is referred to as pre-convolution data.is a schematic diagram showing the pre-convolution data.

6 3 FIG. 3 FIG. The convolution layer processing unitperforms a convolution layer process with a filter size of 1×1 on the pre-convolution data, thereby to generate 3 dimensional data in which the number of elements (A) in predetermined one dimension (C dimension) is restored to the original number of elements (C). The size of this 3 dimensional data is H×W×C.is a schematic diagram showing the pre-convolution data and 3 dimensional data (hereinafter referred to as 3 dimensional data Y) obtained after the execution of the convolution layer process with the filter size of 1×1. In, the elements of C dimension are shown vertically aligned for convenience.

before after The elements in the 3 dimensional data can be specified by 3 dimension coordinates. In the pre-convolution data, the values of the elements specified by the H dimension coordinate h, the W dimension coordinate w, and the C dimension coordinate c are expressed as (h, w, c). Similarly, in the 3 dimensional data Y, the values of the elements specified by the H dimension coordinate h, the W dimension coordinate w, and the C dimension coordinate c are expressed as (h, W, C).

6 The size 1×1 filter values used by the convolution layer processing unitare defined as C sets of A filter values as 1 set. That is, A×C filter values are determined. The A×C filter values are determined in advance by learning.

The 0th set of A filter values is used to calculate the value of each element of the 0th channel in the 3 dimensional data Y. Similarly, the i-th set of A filter values (i is an integer such that 0≤i≤C−1) is used to calculate the value of each element of the i-th channel in the 3 dimensional data Y.

(0, 0) (0, 1) (0, A-1) alter 6 For example, it is assumed that the 0th set of A filter values is a, a, . . . , a, in order from the 0th. In this case, the convolution layer processing unitcalculates the value of (0, 0, 0)by the following formula (1).

6 The convolution layer processing unitalso obtains the values of the other elements of the 0th channel in the 3 dimensional data Y using the 0th set of A filter values by the similar calculation.

(i, 0) i, 1) (i, A−1) after 6 It is assumed that the i-th set of A filter values is a, a, . . . , a, in order from the 0th. In this case, the convolution layer processing unitcalculates the value of (0, 0, i)by the following formula (2).

6 The convolution layer processing unitalso obtains the values of the other elements of the i-th channel in the 3 dimensional data Y using the i-th set of A filter values by the similar calculation.

6 6 6 6 The convolution layer processing unituses the above calculation to calculate the value of each element of the 0th channel, the value of each element of the 1st channel, . . . , the value of each element of the C-1st channel in the 3 dimensional data Y. Then, the convolution layer processing unitperforms the same process for all positions in the plane consisting of the H dimension and W dimension in the pre-convolution data. In other words, the convolution layer processing unitcalculates the values of all elements in the 3 dimensional data Y. Then, the convolution layer processing unitderives the 3 dimensional data Y. As a result, 3 dimensional data Y whose size is H×W×C is obtained.

7 6 7 7 1 7 7 7 1 The Sigmoid layer processing unitapplies a sigmoid function to the individual elements in the 3 dimensional data Y derived by the convolution layer processing unit. As a result, the value of each element in the 3 dimensional data changes to a value in the range of 0 to 1. The size of the 3 dimensional data is not changed by the processing by the Sigmoid layer processing unit. The 3 dimensional data after the processing by the Sigmoid layer processing unitis denoted by the symbol Y′. The neural network processing devicemay not include the Sigmoid layer processing unit, and the above processing by the Sigmoid layer processing unitmay be omitted. Also, instead of the Sigmoid layer processing unit, a component that applies a function other than the sigmoid function to individual elements in the 3 dimensional data Y may be included in the neural network processing device.

6 104 after after 3 FIG. 8 FIG. The A×C filter values (value of filter of filter size 1×1) used by the convolution layer processing unitcan also be referred to as weights. Then, it can be said that the calculation of C values from (0, 0, 0)to (0, 0, C−1)(see right side of) is the same as the process to obtain 1 dimensional data with C elements by obtaining the values of C elements in the second FC layer in the SE block (see step Sin).

7 106 8 FIG. Therefore, the 3 dimensional data Y′ obtained by the Sigmoid layer processing unitis the same as the 3 dimensional data X′ obtained by the copy process performed in the Scale layer of the SE block (see step Sin).

Therefore, each element of each channel in the 3 dimensional data Y′ is a coefficient that represents the degree of importance of the channel corresponding to that element (the channel of the 3 dimensional data U). For example, the values of the H×W elements in the 0th channel of the 3 dimensional data Y′ are common, and are the coefficients representing the degree of importance of the 0th channel of the 3 dimensional data U.

8 7 8 The Scale layer processing unitgenerates 3 dimensional data by performing an element-wise product operation of the first input 3 dimensional data U and the 3 dimensional data Y′ obtained by the Sigmoid layer processing unit. The Scale layer processing unitoutputs the 3 dimensional data.

4 FIG. 7 8 8 is a schematic diagram showing calculation of the element-wise product of the 3 dimensional data U and the 3 dimensional data Y′ obtained by the Sigmoid layer processing unit. The sizes of both the 3 dimensional data U and 3 dimensional data Y′ are H×W×C, and are common. Furthermore, the elements in the 3 dimensional data U and the elements in the 3 dimensional data Y′ can both be specified by 3 dimension coordinates. Therefore, it is possible to associate elements in the 3 dimensional data U and elements in the 3 dimensional data Y′ that share the same 3 dimension coordinates. As a result, the elements in the 3 dimensional data U and the elements in the 3 dimensional data Y′ are associated one-to-one. The Scale layer processing unitthe generates 3 dimensional data by calculating the product of the values of elements for each pair of elements to be associated (in other words, by performing the element-wise product operation). The 3 dimensional data generated by the Scale layer processing unitcan be said to be the data obtained by multiplying the multiple elements for each individual channel of the 3 dimensional data U by the coefficient corresponding to the channel (coefficient representing the degree of importance).

2 3 4 5 6 7 8 2 3 4 5 6 7 8 The 1 dimensional data generation unit, the FC layer processing unit, the ReLU layer processing unit, the copy unit, the convolution layer processing unit, The Sigmoid layer processing unitand the Scale layer processing unitare realized, for example, by a CPU (Central Processing Unit) of a computer operating according to a neural network processing program. For example, the CPU may read the neural network processing program from a program storage medium such as a program storage device of the computer, and operate as the 1 dimensional data generation unit, the FC layer processing unit, the ReLU layer processing unit, the copy unit, the convolution layer processing unit, The Sigmoid layer processing unitand the Scale layer processing unitaccording to the neural network processing program.

2 3 4 5 6 7 8 Alternatively, the 1 dimensional data generation unit, the FC layer processing unit, the ReLU layer processing unit, the copy unit, the convolution layer processing unit, the Sigmoid layer processing unit, and the Scale layer processing unitmay each be realized by separate hardware.

5 FIG. 9 FIG. Next, the processing flow will be described.is a flowchart showing an example of the processing flow of the example embodiment of the present invention. In the following explanation, the case in which the input multidimensional data is the aforementioned 3 dimensional data U (see) will be used as an example. In addition, detailed explanations of matters already explained will be omitted.

2 2 1 When the 1 dimensional data generation unitinputs of the 3 dimensional data U, the 1 dimensional data generation unitgenerates the 1 dimensional data whose number of elements is the number of elements of the C dimension (C) in the 3 dimensional data U, by setting the number of elements of each dimension other than the C dimension (i.e., the H dimension and the W dimension) to 1 (step S).

3 1 2 Next, the FC layer processing unitreduces the number of elements in the 1 dimensional data obtained in step Sby performing FC layer processing (step S). It is assumed that the number of elements after the reduction is A. A<C.

4 2 3 4 Next, the ReLU layer processing unitchanges the values of the elements in the 1 dimensional data obtained in step Sthat have negative values to 0 (step S). At this time, the ReLU layer processing unitdoes not change the value of elements whose value is equal to or greater than 0.

5 3 4 2 FIG. Next, the copy unitgenerates the pre-convolution data (see) whose size is H×W×A by copying the 1 dimensional data obtained in step SH×W times (Step S).

6 5 Next, the convolution layer processing unitperforms convolution layer processing with a filter size of 1×1 on the pre-convolution data, thereby to generate 3 dimensional data Y in which the number of elements of the C dimension is restored to the original number of elements (C) (Step S).

7 6 Next, the Sigmoid layer processing unitderives the 3 dimensional data Y′ by applying a sigmoid function to the individual elements in the 3 dimensional data Y (step S). The value of each element included in the 3 dimensional data Y′ is in the range of 0 to 1.

8 7 Next, the Scale layer processing unitgenerates 3 dimensional data by performing the element-wise product operation of the first input 3 dimensional data U and the 3 dimensional data Y′, and outputs the 3 dimensional data (Step S).

7 106 8 8 FIG. Next, the effect of the present example embodiment is explained. As mentioned above, the 3 dimensional data Y′ obtained by the Sigmoid layer processing unitis the same as the 3 dimensional data X′ obtained by the copy process performed in the Scale layer of the SE block (see step Sin). Therefore, in the present example embodiment, the same result as in the SE block is obtained by the element-wise product operation in the Scale layer processing unit. Therefore, according to the neural network processing device of the present example embodiment, the same CNN accuracy as the SE block can be obtained.

5 4 In the present example embodiment, the copy unitcopies the 1 dimensional data derived by the ReLU layer processing unitmultiple times. The size of this 1 dimensional data is 1×1×A. Also, A<C. Therefore, the overhead of the copy process in the present example embodiment is smaller than the overhead of the copy process in the SE block, which copies 1 dimensional data whose size is 1×1×C multiple times. Therefore, according to the present example embodiment, the operation can be performed at a faster speed than in the SE block.

In other words, according to the present example embodiment, it is possible to obtain the same CNN accuracy as the SE block and perform operations faster than the SE block. As a result, processing time can be reduced.

6 104 6 8 FIG. Note that the convolution layer processing unitperforms convolution layer process with a filter size of 1×1 on the pre-convolution data, therefore the amount of operations required to obtain the 3 dimensional data Y is larger than the amount of operations of the second FC layer in the SE block (see step Sin). However, the speed of the convolution layer process is fast. Therefore, even if the amount of operations performed by the convolution layer processing unitin obtaining the 3 dimensional data Y is large, the effect on the processing speed (processing time) is small, and as a result, the processing speed can be faster than in the SE block.

1 7 5 FIG. In a neural network, a block is a batch of multiple layers which are basic components, and the block is applied multiple times. Then, steps S-S(see) of the present example embodiment can be considered as one block. This block is denoted as block P. The block P can be applied in multiple location in the neural network process. Here, the degree of the effect of block P (the effect of the present example embodiment) varies depending on the size of the input multidimensional data. Therefore, in the neural network process, block P may be applied to the location where the effect of block P is large. In the neural network process, the location where the effect is large may be specified in advance by experimentation or other means.

6 FIG. 1000 1001 1002 1003 1004 is a schematic block diagram showing an example of computer configuration of the neural network processing device of the example embodiment of the present invention. The computerincludes a CPU, a main memory, an auxiliary memory, and an interface.

1 1000 1 1003 1001 1003 1002 The neural network processing deviceof the example embodiment of the present invention is realized by a computer. The operation of the neural network processing deviceis stored in the auxiliary memoryin the form of a neural network processing program. The CPUreads the neural network processing program from auxiliary memoryand expands it to the main memory, and executes the process described in the above example embodiment.

1003 1004 1000 1000 1002 The auxiliary memoryis an example of a non-transitory tangible medium. Other examples of non-transitory tangible media include magnetic disks connected via interface, magneto-optical disks, CD-ROM (Compact Disk Read Only Memory), DVD-ROM (Digital Versatile Disk Read Only Memory), semiconductor memory, etc. When the program is delivered to the computerthrough a communication line, the computermay expand the program in the main memoryand execute the process described in the above example embodiment according to the program.

Some or all of the components may be realized by general-purpose or dedicated circuitry, processor, or a combination of these. These may comprise a single chip or multiple chips connected via a bus. Some or all of the components may be realized by a combination of the above-mentioned circuitry, etc. and a program.

When some or all of components is realized by multiple information processing devices, circuits, etc., the multiple information processing devices, circuits, etc. may be centrally located or distributed. For example, the information processing devices and circuits may be realized as a client-and-server system, a cloud computing system, etc., each of which is connected via a communication network.

7 FIG. 92 93 95 96 98 The following is an overview of the invention.is a block diagram showing an overview of the neural network processing device of the present invention. The neural network processing device of the present invention includes 1 dimensional data generation means, number of elements reducing means, copy means, convolution layer processing means, and element-wise product operation means.

92 2 The 1 dimensional data generation means(e.g., the 1 dimensional data generation unit) generates, by setting the number of elements of each dimension (e.g., the H dimension and the W dimension) other than predetermined one dimension (e.g., the C dimension (dimension of channel)) to 1 based on multidimensional data corresponding to one input data (e.g., the 3 dimensional data U), 1 dimensional data whose number of elements is the same as the number of elements of the predetermined one dimension in the multidimensional data (e.g., C).

93 3 The number of elements reducing means(e.g., the FC layer processing unit) reduces the number of elements included in the 1 dimensional data.

95 5 The copy means(e.g., the copy unit) copies the 1 dimensional data whose number of elements has been reduced multiple times, thereby to generate multidimensional data (e.g., pre-convolution data) in which the number of elements of each dimension (e.g., the H dimension and the W dimension) other than the predetermined one dimension is restored to original number of elements (e.g., H and W).

96 6 95 95 The convolution layer processing means(e.g., convolution layer processing unit) performs a convolution layer process with a filter size of 1×1 on the multidimensional data generated by the copy means, thereby to generate multidimensional data in which the number of element of the predetermined one dimension (e.g., the C dimension) in multidimensional data generated by the copy meansis restored to original number of elements (e.g., C).

98 The element-wise product operation meansperforms an element-wise product operation, based on the multidimensional data corresponding to one input data (e.g., the 3 dimensional data U) and the multidimensional data generated by the convolution layer processing means.

According to such a configuration, it is possible to obtain the same CNN accuracy as the SE block and perform operations faster than the SE block.

7 96 98 The neural network processing device may include sigmoid function processing means (e.g., the Sigmoid layer processing unit) for applying a sigmoid function to each element in the multidimensional data generated by the convolution layer processing means, and the element-wise product operation meansmay perform the element-wise product operation of the multidimensional data corresponding to one input data and the multidimensional data after the sigmoid function is applied to each element.

4 95 The neural network processing device may include change means (e.g., the ReLU layer processing unit) for changing values of the elements in the 1 dimensional data whose number of elements has been reduced that have negative values to 0, and the copy meansmay copy the 1 dimensional data after processing by the change means multiple times.

The multidimensional data corresponding to one input data may be 3 dimensional data, and the predetermined one dimension may be a dimension of channel.

Although the present invention has been described above with reference to example embodiment, the present invention is not limited to the above example embodiment. Various changes may be made to the structure and details of the present invention, that may be understood by those skilled in the art within the scope of the present invention.

The invention is suitable for a neural network processing device that performs a process of a neural network.

1 Neural network processing device 2 1 dimensional data generation unit 3 FC layer processing unit 4 ReLU layer processing unit 5 Copy unit 6 Convolution layer processing unit 7 Sigmoid layer processing unit 8 Scale Layer processing unit

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

May 25, 2020

Publication Date

July 7, 2026

Inventors

Seiya Shibata

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “Multidimensional data processing device, method, and computer-readable recording medium” (US-12675673-B2). https://patentable.app/patents/US-12675673-B2

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.