A processing device comprises a first set of processors comprising a first processor and a second processor, each of which comprises at least one controllable port, a first memory operably coupled to the first set of processors, at least one forward data line configured for one-way transmission of data in a forward direction between the first set of processors, and at least one backward data line configured for one-way transmission of data in a backward direction between the first set of processors. wherein the first set of processors are operably coupled in series via the at least one forward data line and the at least one backward data line.
Legal claims defining the scope of protection, as filed with the USPTO.
a plurality of neural cores, each of which comprises a plurality of controllable ports; at least one forward data line configured for one-way transmission of data in a forward direction between the plurality of neural cores; and at least one backward data line configured for one-way transmission of data in a backward direction between the plurality of neural cores, wherein a data path for the plurality of neural cores is determined among a plurality of candidate data paths according to a type of computation performed by the at least one of the plurality of neural cores, the determined data path is configured on the at least one forward data line and the at least one backward data line by controlling controllable ports of the plurality of neural cores, and the controlling controllable ports comprises turning on or off at least one controllable port of the controllable ports of the plurality of neural cores, in association with determining the determined data path, turning on at least one controllable port corresponding to the determined data path and turning off at least one controllable port corresponding to a non-selected data path among the plurality of candidate data paths. wherein controlling the controllable ports further comprises: . A processor comprising:
claim 1 . The processor of, wherein the plurality of neural cores is operably coupled in series via the at least one forward data line and the at least one backward data line.
claim 1 wherein a forward data path is configured by turning on a controllable port on the at least one forward data line between the first neural core and the second neural core and data is transmitted via the forward data path from the first neural core to the second neural core. . The processor of, wherein the plurality of neural cores comprises a first neural core and a second neural core,
claim 3 . The processor of, wherein a backward data path is configured by turning on a controllable port on the at least one backward data line between the first neural core and the second neural core, and data is transmitted via the backward data path from the second neural core to the first neural core.
claim 1 wherein the plurality of neural cores comprises a first neural core and a second neural core, wherein the determined data path comprises a first and second forward sub-paths configured on the at least one forward data line and a first and second backward sub-paths configured on the at least one backward data line, and wherein the first forward sub-path is from the first memory to the first neural core, the second forward sub-path is from the first neural core to the second neural core, the first backward sub-path is from the second neural core to the first neural core, and the second backward sub-path is from the first neural core to the first memory. . The processor of, further comprising a first memory operably coupled to the plurality of neural cores,
claim 5 the third forward sub-path is from the second neural core to a second memory, and the third backward sub-path is from the second memory to the second neural core, and wherein the second memory is the same as or different from the first memory. . The processor of, wherein the determined data path further comprises a third forward sub-path configured on the at least one forward data line and a third backward sub-path configured on the at least one backward data line,
claim 1 wherein the plurality of neural cores comprises a first neural core and a second neural core, wherein a first data path is configured by controlling controllable ports of the plurality of neural cores, the first data path comprises a forward sub-path configured on the at least one forward data line and a backward sub-path configured on the at least one backward data line, the forward sub-path of the first data path is from the first memory to the first neural core, and the backward sub-path of the first data path is from the first neural core to the first memory, wherein a second data path is configured by controlling controllable ports of the plurality of neural cores, the second data path comprises a forward sub-path configured on the at least one forward data line and a backward sub-path configured on the at least one backward data line, the backward sub-path of the second data path is from a second memory to the second neural core, and the forward sub-path of the second data path is from the second neural core to the second memory, and wherein the second memory is the same as or different from the first memory. . The processor of, further comprising a first memory operably coupled to the plurality of neural cores,
claim 1 wherein the determined data path comprises first and second forward sub-paths configured on the at least one forward data line, the first forward sub-path is from the first memory to the first neural core, and the second forward sub-path is from the first neural core to a second memory, and wherein the second memory is the same as or different from the first memory. . The processor of, wherein the plurality of neural cores comprises a first neural core and a second neural core,
claim 1 some of the plurality of forward data lines are turned off by controlling controllable ports of the plurality of neural cores. . The processor of, wherein the at least one forward data line comprises a plurality of forward data lines, and
claim 1 wherein the first set of neural cores comprises a first neural core and a second neural core, the second set of neural cores comprises a third neural core and a fourth neural core, wherein the first set of neural cores are connected in series with each other via a forward data line and a backward data line between the first set of neural cores, wherein the second set of neural cores are connected in series with each other via a forward data line and a backward data line between the second set of neural cores. . The processor of, wherein the plurality of neural cores comprises a first set of neural cores and a second set of neural cores,
claim 10 an interconnection through which data are moved, wherein the first memory, the first set of neural cores, and the second set of neural cores are connected to the interconnection, and data are moved among the first memory, the first set of neural cores, and the second set of neural cores via the interconnection. . The processor of, further comprising a first memory operably coupled to the plurality of neural cores, and
claim 1 wherein the processor further comprising: at least one first connection line connected to the at least one forward data line and the first neural core; at least one second connection line connected to the at least one backward data line and the first neural core; at least one third connection line connected to the at least one forward data line and the second neural core; and at least one fourth connection line connected to the at least one backward data line and the second neural core. . The processor of, wherein the plurality of neural cores comprises a first neural core and a second neural core, and
claim 1 wherein the plurality of neural cores are included in a system-on-chip, and the first memory comprises an off-chip memory external to the system-on-chip. . The processor of, further comprising a first memory operably coupled to the plurality of neural cores, and
claim 1 wherein the first set of neural cores comprises a first neural core and a second neural core, the second set of neural cores comprises a third neural core and a fourth neural core, and a shared memory shared by the first set of neural cores and the second set of neural cores. . The processor of, wherein the plurality of neural cores comprises a first set of neural cores and a second set of neural cores,
claim 1 . The processor of, wherein controllable ports of the plurality of neural cores are implemented by software or firmware.
a plurality of neural cores comprising a first neural core, a second neural core connected in series with the first neural core in a forward direction, and a third neural core connected in series with the second neural core in the forward direction, each neural core comprising a plurality of controllable ports; and a first memory connected in series with the first neural core in a backward direction, a data path for the plurality of neural cores is determined among a plurality of candidate data paths according to a type of a computation performed by the at least one of the plurality of neural cores, the determined data path is configured on the at least one forward data line and the at least one backward data line by controlling controllable ports of the plurality of neural cores, wherein the determined data path comprises a data movement path in the forward direction and a data movement path in the backward direction, the determined data path is configured even when at least one of the first neural core, the second neural core, or the third neural core is inoperative, and the controlling controllable ports comprises turning on or off at least one controllable port of the controllable ports of the plurality of neural cores, in association with determining the determined data path, turning on at least one controllable port corresponding to the determined data path and turning off at least one controllable port corresponding to a non-selected data path among the plurality of candidate data paths. wherein controlling the controllable ports further comprises: . A processor comprising:
claim 16 the first neural core is provided with first data from the first memory and generates second data by computing the first data, the second neural core is provided with the second data from the first neural core and provides the second data to the third neural core, and the third neural core is provided with the second data from the second neural core and generates third data by computing the second data. . The processor of, wherein if the first neural core and the third neural core are operative and the second neural core is inoperative,
claim 16 all of the plurality of lines are turned on or some of the plurality of lines are turned off according to a bandwidth of data provided from the first memory. . The processor of, wherein data lines connecting the first neural core to the third neural core comprise a plurality of lines, and
claim 16 a fourth neural core that is not directly connected to the first neural core; a fifth neural core connected in series with the fourth neural core in the forward direction; and an interconnection configured to perform data exchange between the first neural core and the fourth neural core. . The processor of, further comprising:
receiving a task including information of the data paths for the plurality of neural cores including a first neural core and a second neural core, and determining a data path among the plurality of data paths configured on the at least one forward data line and the at least one backward data line by controlling controllable ports of the plurality of neural cores according to a type of a computation performed by the at least one of the plurality of neural cores, and wherein determining a data path comprises: configuring a forward data path by turning on a controllable port of the first neural core and a controllable port of the second neural core on the at least one forward data line between the first neural core and the second neural core and data is transmitted via the forward data path from the first neural core to the second neural core, and configuring a backward data path by turning on a controllable port of the first neural core and a controllable port of the second neural core on the at least one backward data line between the first neural core and the second neural core and data is transmitted via the backward data path from the second neural core to the first neural core, wherein controlling the controllable ports further comprises: in association with determining the determined data path, turning on at least one controllable port corresponding to the determined data path and turning off at least one controllable port corresponding to a non-selected data path among the plurality of candidate data paths. . A method, performed by a processor, for configuring a plurality of data paths of a plurality of neural cores, comprising:
Complete technical specification and implementation details from the patent document.
This application is a continuation of U.S. application Ser. No. 18/447,226, filed on Aug. 9, 2023, which claims priority under 35 U.S.C § 119 to Korean Patent Application No. 10-2022-0174933 filed on Dec. 14, 2022, in the Korean Intellectual Property Office, the entire contents of which is hereby incorporated by reference.
The disclosure relates to neural processors. More particularly, the disclosure relates to neural processors capable of reconfiguring data paths of neural cores.
For the last few years, artificial intelligence technology has been the core technology of the Fourth Industrial Revolution and the subject of discussion as the most promising technology worldwide. The biggest problem with artificial intelligence technology is computing performance. For artificial intelligence technology to realize a level of human learning ability, reasoning ability, perceptual ability, natural language implementation ability, etc., it is of the utmost importance to process a large amount of data quickly.
The central processing unit (CPU) or graphics processing unit (GPU) of off-the-shelf computers was used to implement deep-learning training and inference in early artificial intelligence, but these components had limitations in their ability to perform the tasks of deep-learning training and inference with high workloads. Thus, neural processing units (NPUs) that are structurally specialized for deep learning tasks have received a lot of attention. These neural processing units have a plurality of computation devices therein, and each computation device operates in parallel and thereby enhance computation efficiency.
Recently, in order to maximize computation efficiency, the trend is to gradually increase the number of cores within a computation device. However, if the number of cores increases, data paths must be assigned newly by the increased number of cores. This affects very disadvantageously in terms of scalability, and thus, there exist drawbacks of not only being hard to miniaturize chips but also increasing the complexity of design.
In addition, if data paths are predetermined, there may be disadvantages that power can be wasted unnecessarily since there may arise a case in which more cores than actually needed for computation must be used, and the efficiency of a computation device may be reduced as the most efficient data path cannot be dynamically determined depending on the type of computations performed by the computation device.
Aspects of the disclosure provide a neural processor capable of reconfiguring data paths.
Aspects of the disclosure provide a neural processor that has relatively high scalability.
Aspects of the disclosure provide a neural processor that has relatively low power consumption.
Aspects of the disclosure provide a neural processor that configures optimized data paths according to data flows.
According to some aspects, a processing device may comprise: a first set of processors comprising a first processor and a second processor, each of which comprises at least one controllable port; a first memory operably coupled to the first set of processors; at least one forward data line configured for one-way transmission of data in a forward direction between the first set of processors; and at least one backward data line configured for one-way transmission of data in a backward direction between the first set of processors, and wherein the first set of processors are operably coupled in series via the at least one forward data line and the at least one backward data line.
According to some aspects, the data paths are configured by controlling controllable ports of the first set of processors.
According to some aspects, a forward data path is configured by turning on a controllable port on the at least one forward data line between the first processor and the second processor and data is transmitted via the forward data path from the first processor to the second processor.
According to some aspects, a backward data path is configured by turning on a controllable port on the at least one backward data line between the first processor the second processor and data is transmitted via the backward data path from the second processor to the first processor.
According to some aspects, a data path is configured by controlling controllable ports of the first set of processors, the data path comprises a first and second forward sub-paths configured on the at least one forward data line and a first and second backward sub-paths configured on the at least one backward data line, the first forward sub-path is from the first memory to the first processor, the second forward sub-path is from the first processor to the second processor, the first backward sub-path is from the second processor to the first processor, and the second backward sub-path is from the first processor to the first memory.
According to some aspects, the data path further comprises a third forward sub-path configured on the at least one forward data line and a third backward sub-path configured on the at least one backward data line, the third forward sub-path is from the second processor to a second memory, the third backward sub-path is from the second memory to the second processor, and the second memory is the same as or different from the first memory.
According to some aspects, a first data path is configured by controlling controllable ports of the first set of processors, the first data path comprises a forward sub-path configured on the at least one forward data line and a backward sub-path configured on the at least one backward data line, the forward sub-path of the first data path is from the first memory to the first processor, and the backward sub-path of the first data path is from the first processor to the first memory, and a second data path is configured by controlling controllable ports of the first set of processors, the second data path comprises a forward sub-path configured on the at least one forward data line and a backward sub-path configured on the at least one backward data line, the backward sub-path of the second data path is from a second memory to the second processor, and the forward sub-path of the second data path is from the second processor to the second memory, wherein the second memory is the same as or different from the first memory.
According to some aspects, a data path is configured by controlling controllable ports of the first set of processors, the data path comprises first and second forward sub-paths configured on the at least one forward data line, the first forward sub-path is from the first memory to the first processor, and the second forward sub-path is from the first processor to a second memory, wherein the second memory is the same as or different from the first memory.
According to some aspects, the at least one forward data line comprises a plurality of forward data lines, and some of the plurality of forward data lines are turned off by controlling controllable ports of the first set of processors.
According to some aspects, the processing device may further comprise a second set of processors comprising a third processor and a fourth processor, wherein the second set of processors are connected in series with each other via a forward data line and a backward data line between the second set of processors.
According to some aspects, the processing device may further comprise an interconnection through which data are moved, wherein the memory, the first set of processors, and the second set of processors are connected to the interconnection, and data are moved between the memory, the first set of processors, and the second set of processors via the interconnection.
According to some aspects, the processing device may further comprise at least one first connection line connected to the at least one forward data line and the first processor; at least one second connection line connected to the at least one backward data line and the first processor; at least one third connection line connected to the at least one forward data line and the second processor; and at least one fourth connection line connected to the at least one backward data line and the second processor.
According to some aspects, the first set of processors are included in a system-on-chip, and the first memory comprises an off-chip memory external to the system-on-chip.
According to some aspects, the processing device may further comprise a shared memory shared by the first set of processors and the second set of processors, wherein the memory comprises the shared memory.
According to some aspects, controllable ports of the first set of processors are implemented by software or firmware.
According to some aspects, a processing device may comprise a first processor; a second processor connected in series with the first processor in a forward direction; a third processor connected in series with the second processor in the forward direction; and a first memory connected in series with the first processor in a backward direction, wherein data paths for the first core, the second core, the third core, and the first memory are configured, wherein the data paths comprise a data movement path in the forward direction and a data movement path in the backward direction, and the data paths are configured even when at least one of the first processor, the second processor, or the third processor is inoperative.
According to some aspects, wherein if the first processor and the third processor are operative and the second processor is inoperative, the first processor is provided with first data from the first memory and generates second data by computing the first data, the second processor is provided with the second data from the first processor and provides the second data to the third processor, and the third processor is provided with the second data from the second processor and generates third data by computing the second data.
According to some aspects, data lines connecting the first core to the third core comprise a plurality of lines, and all of the plurality of lines are turned on or some of the plurality of lines are turned off according to a bandwidth of data provided from the first memory.
According to some aspects, the processing device may further comprise a fourth processor that is not directly connected to the first processor; a fifth processor connected in series with the fourth processor in the forward direction; and an interconnection configured to perform data exchange between the first processor and the fourth processor.
According to some aspects, a processing device may comprise a first processor; a second processor connected in series with the first processor in a forward direction; a third processor connected in series with the second processor in the forward direction; and a first memory connected in series with the first processor in a backward direction, wherein data paths for the first processor, the second processor, the third processor, and the first memory are configured in real time.
Aspects of the disclosure are not limited to those mentioned above and other objects and advantages of the disclosure that have not been mentioned can be understood by the following description and will be more clearly understood according to embodiments of the disclosure. In addition, it will be readily understood that the objects and advantages of the disclosure can be realized by the means and combinations thereof set forth in the claims.
Even if the number of neural cores is changed, the neural processor of the disclosure is relatively simple to change the design resulting therefrom and thus has relatively high scalability.
The neural processor of the disclosure can reconfigure data paths relatively easily, and can increase computation efficiency by appropriately changing the data paths as necessary.
The neural processor of the disclosure can minimize power consumption by turning off unused neural cores or turning off some of the unused data lines.
In addition to the foregoing, the specific effects of the disclosure will be described together while elucidating the specific details for carrying out the embodiments below.
The terms or words used in the disclosure and the claims should not be construed as limited to their ordinary or lexical meanings. They should be construed as the meaning and concept in line with the technical idea of the disclosure based on the principle that the inventor can define the concept of terms or words in order to describe his/her own embodiments in the best possible way. Further, since the embodiment described herein and the configurations illustrated in the drawings are merely one embodiment in which the disclosure is realized and do not represent all the technical ideas of the disclosure, it should be understood that there may be various equivalents, variations, and applicable examples that can replace them at the time of filing this application.
Although terms such as first, second, A, B, etc. used in the description and the claims may be used to describe various components, the components should not be limited by these terms. These terms are used only for the purpose of distinguishing one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may be referred to as a first component, without departing from the scope of the disclosure. The term ‘and/or’includes a combination of a plurality of related listed items or any item of the plurality of related listed items.
The terms used in the description and the claims are merely used to describe particular embodiments and are not intended to limit the disclosure. Singular expressions include plural expressions unless the context explicitly indicates otherwise. In the application, terms such as “comprise,” “have,” “include”, “contain,” etc. should be understood as not precluding the possibility of existence or addition of features, numbers, steps, operations, components, parts, or combinations thereof described herein. Terms such as a “circuit” or “circuitry”, refers to a circuit in hardware but may also refer to a circuit in software.
Unless otherwise defined, the phrases “A, B, or C,” “at least one of A, B, or C,” or “at least one of A, B, and C” may refer to only A, only B, only C, both A and B, both A and C, both B and C, all of A, B, and C, or any combination thereof.
Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which the disclosure pertains.
Terms such as those defined in commonly used dictionaries should be construed as having a meaning consistent with the meaning in the context of the relevant art, and are not to be construed in an ideal or excessively formal sense unless explicitly defined in the disclosure.
In addition, each configuration, procedure, process, method, or the like included in each embodiment of the disclosure may be shared to the extent that they are not technically contradictory to each other.
1 32 FIGS.to Hereinafter, a neural processing device in accordance with some embodiments of the disclosure will be described with reference to.
1 FIG. is a block diagram illustrating a neural processing system in accordance with some embodiments of the disclosure.
1 FIG. 1 2 3 Referring to, a neural processing system NPS in accordance with some embodiments may include a first neural processing device, a second neural processing device, and an external interface.
1 1 The first neural processing devicemay be a device that performs calculations using an artificial neural network. The first neural processing devicemay be, for example, a device specialized in performing tasks of deep learning computations. However, the embodiment is not limited thereto.
2 1 1 2 3 The second neural processing devicemay be a device having the same or similar configuration as the first neural processing device. The first neural processing deviceand the second neural processing devicemay be connected to each other via the external interfaceand share data and control signals.
1 FIG. 3 Althoughshows two neural processing devices, the neural processing system NPS in accordance with some embodiments is not limited thereto. In some embodiments, in a neural processing system NPS, three or more neural processing devices may be connected to each other via the external interface. Also, conversely, a neural processing system NPS in accordance with some embodiments may include only one neural processing device.
1 2 1 2 1 2 In this case, each of the first neural processing deviceand the second neural processing devicemay be a processing device other than the neural processing device. In some embodiments, each of the first neural processing deviceand the second neural processing devicemay be a graphics processing unit (GPU), a central processing unit (CPU), and other types of processing units as well. In the following, the first neural processing deviceand the second neural processing devicewill be described as neural processing devices for convenience.
2 FIG. 1 FIG. is a block diagram for illustrating the neural processing device of.
2 FIG. 1 10 20 30 40 50 60 70 80 Referring to, a first neural processing devicemay include a neural core SoC, a CPU, an off-chip memory, a first non-volatile memory interface, a first volatile memory interface, a second non-volatile memory interface, a second volatile memory interfaceand a control interface (CIF).
10 10 10 The neural core SoCmay be a system on a chip device. The neural core SoCcan be an artificial intelligence computation device and may be an accelerator. The neural core SoCmay be, for example, any one of a graphics processing unit (GPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). However, the embodiment is not limited thereto.
10 3 10 31 32 40 50 The neural core SoCmay exchange data with other external computation devices via the external interface. Further, the neural core SoCmay be connected to the non-volatile memoryand the volatile memoryvia the first non-volatile memory interfaceand the first volatile memory interface, respectively.
20 1 20 10 The CPUmay be a control device that controls the system of the first neural processing deviceand executes program computations. The CPUis a general-purpose computation device and may have low efficiency in performing simple parallel computations that are frequently used in deep learning. Accordingly, there can be high efficiency by performing computations in deep learning inference and training tasks by the neural core SoC.
20 3 20 31 32 60 70 The CPUmay exchange data with other external computation units via the external interface. Further, the CPUmay be connected to the non-volatile memoryand the volatile memoryvia the second non-volatile memory interfaceand the second volatile memory interface, respectively.
20 10 20 10 10 20 The CPUmay also transfer tasks to the neural core SoCvia commands. In some embodiments, the CPUmay be a kind of host that gives instructions to the neural core SoC. In some embodiments, the neural core SoCcan efficiently perform parallel computation tasks such as deep learning tasks according to the instructions of the CPU.
30 10 30 31 32 The off-chip memorymay be a memory disposed outside the chip of the neural core SoC. The off-chip memorymay include a non-volatile memoryand a volatile memory.
31 31 The non-volatile memorymay be a memory that continuously retains stored information even if electric power is not supplied. The non-volatile memorymay include, for example, at least one of Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Alterable ROM (EAROM), Erasable Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM) (e.g., NAND Flash memory, NOR Flash memory), Ultra-Violet Erasable Programmable Read-Only Memory (UVEPROM), Ferroelectric Random-Access Memory (FeRAM), Magnetoresistive Random-Access Memory (MRAM), Phase-change Random-Access Memory (PRAM), silicon-oxide-nitride-oxide-silicon (SONOS), Resistive Random-Access Memory (RRAM), Nanotube Random-Access Memory (NRAM), magnetic computer storage devices (e.g., hard disks, diskette drives, magnetic tapes), optical disc drives, or 3D XPoint memory. However, the embodiment is not limited thereto.
32 31 32 The volatile memorymay be a memory that continuously requires electric power to retain stored information, unlike the non-volatile memory. The volatile memorymay include, for example, at least one of Dynamic Random-Access Memory (DRAM), Static Random-Access Memory (SRAM), Synchronous Dynamic Random-Access Memory (SDRAM), or Double Data Rate SDRAM (DDR SDRAM). However, the embodiment is not limited thereto.
40 60 Each of the first non-volatile memory interfaceand the second non-volatile memory interfacemay include, for example, at least one of Parallel Advanced Technology Attachment (PATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), Serial Advanced Technology Attachment (SATA), or PCI Express (PCIe). However, the embodiment is not limited thereto.
50 70 Each of the first volatile memory interfaceand the second volatile memory interfacemay be, for example, at least one of SDR (Single Data Rate), DDR (Double Data Rate), QDR (Quad Data Rate), or XDR (eXtreme Data Rate, Octal Data Rate). However, the embodiment is not limited thereto.
80 20 10 80 20 10 80 The control interfacemay be an interface for transferring control signals between the CPUand the neural core SoC. The control interfacemay transmit commands of the CPUand transmit responses thereto of the neural core SoC. The control interfacemay be, for example, PCIe (PCI Express), but is not limited thereto.
3 FIG. 2 FIG. is a block diagram for illustrating the neural core SoC of.
2 3 FIGS.and 10 1000 2000 3000 4000 5000 7000 6000 Referring to, the neural core SoCmay include at least one neural processor, a shared memory, a direct memory access (DMA), a non-volatile memory controller, a volatile memory controller, a command processor, and a global interconnection.
1000 1000 1000 1000 6000 The neural processormay be a computation device that directly performs computation tasks. If there exist a plurality of neural processors, computation tasks may be assigned to respective neural processors. The respective neural processorsmay be connected to each other via the global interconnection.
2000 1000 2000 1000 2000 30 1000 2000 1000 30 2 FIG. The shared memorymay be a memory shared by multiple neural processors. The shared memorymay store data of neural processors. In addition, the shared memorymay receive data from the off-chip memory, store the data temporarily, and transfer the data to neural processors. The shared memorymay also receive data from the neural processor, store the data temporarily, and transfer the data to the off-chip memoryof.
2000 2000 2000 The shared memorymay be required to be a relatively high-speed memory. Accordingly, the shared memorymay include, for example, an SRAM. However, the embodiment is not limited thereto. That is, the shared memorymay include a DRAM as well.
2000 2000 The shared memorymay be a memory corresponding to the SoC level, i.e., level 2 (L2). Accordingly, the shared memorymay also be defined as an L2 shared memory.
3000 1000 20 3000 1000 20 The DMAmay directly control movements of data without needs for the neural processoror CPUto control the input/output of data. Accordingly, the DMAmay control data movements between memories, thereby minimizing a number of interrupts of the neural processoror CPU.
3000 2000 30 3000 4000 5000 The DMAmay control the data movements between the shared memoryand the off-chip memory. Via the authority of the DMA, the non-volatile memory controllerand the volatile memory controllermay perform movements of data.
4000 31 4000 31 40 4000 4000 The non-volatile memory controllermay control tasks of reading from or writing onto the non-volatile memory. The non-volatile memory controllermay control the non-volatile memoryvia the first non-volatile memory interface. In this case, the non-volatile memory controllermay be referred to as a non-volatile memory controller circuit, but for the sake of convenience, the terms are unified as a non-volatile memory controller. In addition, the non-volatile memory controllermay be implemented as a circuit or circuitry.
5000 32 5000 32 5000 32 50 5000 5000 The volatile memory controllermay control tasks of reading from or writing onto the volatile memory. Further, the volatile memory controllermay perform a refresh task of the volatile memory. The volatile memory controllermay control the volatile memoryvia the first volatile memory interface. Likewise, the volatile memory controllermay be referred to as a volatile memory controller circuit, but for the sake of convenience, the terms are unified as a volatile memory controller. In addition, the volatile memory controllermay be implemented as a circuit or circuitry.
7000 80 7000 20 80 7000 20 1000 7000 1000 The command processormay be connected to the control interface. The command processormay receive control signals from the CPUvia the control interface. The command processormay generate tasks via the control signals received from the CPUand transmit the control signals to neural processors. Further, the command processormay receive completion reports for the tasks from neural processors.
6000 1000 2000 3000 4000 7000 5000 3 6000 6000 1000 2000 3000 4000 5000 7000 3 The global interconnectionmay connect the at least one neural processor, the shared memory, the DMA, the non-volatile memory controller, the command processor, and the volatile memory controllerto one another. In addition, the external interfacemay also be connected to the global interconnection. The global interconnectionmay be a path through which data travel between the at least one neural processor, the shared memory, the DMA, the non-volatile memory controller, the volatile memory controller, the command processor, and the external interface.
6000 1000 7000 The global interconnectionmay transmit not only data but also control signals and signals for synchronization. In the neural processing device in accordance with some embodiments of the disclosure, each neural processormay directly transmit and receive the synchronization signals. Accordingly, latencies due to transmissions of the synchronization signals generated by the command processorcan be minimized.
1000 1000 1000 7000 20 In some embodiments, if there exist a plurality of neural processors, there may be dependencies of individual tasks in which a task of one neural processorneeds to be finished before the next neural processorcan start a new task. The end and start of these individual tasks can be checked via the synchronization signals, and in the conventional techniques, the command processoror the host, i.e., the CPU, was exclusively responsible for both receiving these synchronization signals and instructing the start of a new task.
1000 However, as the number of neural processorsincreases and task dependencies are designed more complicatedly, the number of requests and instructions for this synchronization task can increase exponentially. Therefore, the latency resulting from each request and instruction can greatly reduce the efficiency of tasks.
1000 7000 1000 1000 7000 Therefore, in the neural processing device in accordance with some embodiments of the disclosure, each neural processor, instead of the command processor, may directly transmit some of the synchronization signals to other neural processorsaccording to task dependencies. In this case, several neural processorscan perform the synchronization tasks in parallel as compared with the method managed by the command processor, thereby minimizing the latency due to synchronization.
7000 1000 1000 1000 In addition, the command processorneeds to perform the task scheduling of the neural processorsaccording to a task dependency, and the overhead of such scheduling may increase significantly as the number of neural processorsincreases. Therefore, in the neural processing device in accordance with some embodiments of the disclosure, scheduling tasks are also performed in part by individual neural processors, and hence scheduling burden resulting therefrom can be reduced, thereby improving the performance of the device.
1000 7000 7000 Furthermore, the neural processing device in accordance with some embodiments of the disclosure can carry out monitoring whether a task is completed, an event occurs, a task is delayed, or the like in the neural cores of each neural processor, and can minimize intervention of the command processorand reduce load on the command processor, thereby improving the performance of the device.
7000 7000 Moreover, the neural processing device in accordance with some embodiments of the disclosure can selectively generate completion reports by setting whether to monitor tasks for each task. And the neural processing device in accordance with some embodiments of the disclosure can be configured to modify whether to generate a completion report if a report to the command processoris required. Accordingly, it may be possible to report tasks that require an alert without carrying out monitoring all tasks, and stable monitoring of tasks may be possible while reducing the load on the command processor.
4 FIG. 3 FIG. is a structural diagram for illustrating the global interconnection of.
4 FIG. 6000 6100 6200 6300 Referring to, the global interconnectionmay include a data channel, a control channel, and an L2 sync channel.
6100 6100 1000 2000 3000 4000 5000 3 The data channelmay be a dedicated channel for transmitting data. Through the data channel, the at least one neural processor, the shared memory, the DMA, the non-volatile memory controller, the volatile memory controller, and the external interfacemay exchange data with one another.
6200 6200 1000 2000 3000 4000 5000 7000 3 7000 1000 The control channelmay be a dedicated channel for transmitting control signals. Through the control channel, the at least one neural processor, the shared memory, the DMA, the non-volatile memory controller, the volatile memory controller, the command processor, and the external interfacemay exchange control signals with one another. In particular, the command processormay transmit various control signals to neural processors.
6300 6300 1000 2000 3000 4000 5000 7000 3 The L2 sync channelmay be a dedicated channel for transmitting synchronization signals. Through the L2 sync channel, the at least one neural processor, the shared memory, the DMA, the non-volatile memory controller, the volatile memory controller, the command processor, and the external interfacemay exchange synchronization signals with one another.
6300 6000 6000 The L2 sync channelmay be set as a dedicated channel inside the global interconnection, and thus, may not overlap with other channels and transmit synchronization signals quickly. Accordingly, the neural processing device in accordance with some embodiments does not require new wiring work and may smoothly perform the synchronization task by using the global interconnection.
5 FIG. 1 FIG. is a block diagram for illustrating a flow of control signals of the neural processing device of.
5 FIG. 20 7000 80 Referring to, the CPUmay transfer control signals to the command processorvia the control interface. In this case, a control signal may be a signal instructing execution of each operation, such as a computation task or a data load/store task.
7000 1000 6200 1000 The command processormay receive the control signals and transfer the control signals to at least one neural processorvia the control channel. Each control signal may be stored in the neural processoras each task.
6 FIG. 3 FIG. is a block diagram for schematically illustrating the neural processor of.
6 FIG. 1000 200 300 400 200 300 200 300 Referring to, the neural processormay include a core array CoA, a core global, a task manager, and a memory. In this case, the core globaland the task managermay be referred to respectively as a core global circuit and a task manager circuit. However, for the sake of convenience, the terms are respectively unified as a core global and a task manager. In addition, each of the core globaland the task managermay be implemented as a circuit or circuitry.
100 100 100 1000 100 The core array CoA may include a plurality of neural cores. That is, a particular array of the plurality of neural coresis defined as a core array CoA. The plurality of neural coresmay divide and perform tasks of the neural processor. A number of neural coresmay be 8, for example. However, the embodiment is not limited thereto.
100 200 The neural coresmay receive task information from the core globaland perform tasks according to the task information. In this case, a task may be defined by a control signal, and the task may be any one of memory operations. A memory operation may be, for example, any one of micro-DMA (DMA), LP micro-DMA (Low Priority DMA), store DMA (ST DMA), or a pre-processing task.
100 2000 30 120 120 100 2000 30 20 Specifically, a micro-DMA task may be a task in which the neural coreloads a program or data from the shared memoryor the off-chip memoryto the L0 memory. An LP micro-DMA task may be a load task for a program or data to be used later rather than a current program or data, unlike a general micro-DMA task. Since such a task has a low priority, it can be identified differently from the micro-DMA task. An ST micro-DMA task may be a store task that stores data from the L0 memoryof the neural coreto the shared memoryor the off-chip memory. A pre-processing task may include a task that pre-loads data such as a large number of lookup tables in the CPU.
400 100 1000 400 100 400 100 400 1000 10 400 1000 1000 The memorymay be a memory shared by neural coresin the neural processor. The memorymay store data of neural cores. In addition, the memorymay transfer data to neural cores. On the other hand, the memorymay be a memory shared by neural processorsin the neural core SoC. In this case, the memorymay store data of neural processorsand may transfer the data to neural processors.
400 1000 1000 400 1000 400 2000 2000 1000 400 100 3 FIG. 3 FIG. In some embodiments, the memorymay be an on-chip memory included in the neural processoror an off-chip memory arranged outside the neural processor. For example, the memorymay be an L1 shared memory included in the neural processor, or the memorymay be the shared memoryof. The shared memoryofmay be expressed as an L2 shared memory in another term. The L1 shared memory may be a memory corresponding to the level of the neural processor, i.e., L1 (level 1). The L2 shared memory may be a memory corresponding to the neural processing device, i.e., L2 (level 2). In some embodiments, the L2 shared memory may be shared by the neural processor, and the L1 shared memorymay be shared by the neural core.
100 200 The neural coresmay receive task information from the core globaland perform tasks according to the task information. In this case, a task may be a computation task (or, calculation task) or a task related to a memory operation, and may include information about data paths. A task may be defined by a control signal. The task information is information about a task, and may be information about a type of a task, a form of a task, additional information about a task, and the like.
100 200 The neural coresmay transfer a completion signal indicating completion of execution of a task to the core global.
300 7000 The task managermay receive tasks from a control interconnection CI. In this case, the control interconnection CI may be a generic term for transmission interfaces that transfer tasks from the command processor.
300 200 300 200 7000 The task managermay receive tasks, generate task information, and transmit the task information to the core global. In some embodiments, the task information may include information about data paths. Further, the task managermay receive completion signals via the core global, generate completion reports accordingly, and transmit the completion reports to the command processorvia the control interconnection CI.
200 100 200 100 400 300 The core globalmay be a wire structure connected in hardware within the neural cores. Although not shown, the core globalmay be a structure connecting all of the neural cores, the memory, and the task manager.
200 300 100 100 200 300 The core globalmay receive task information from the task managerand transfer the task information to the neural cores, and may receive completion signals related thereto from the neural cores. Subsequently, the core globalmay transfer the completion signals to the task manager.
7 FIG. is a diagram for illustrating data lines and connection lines connecting neural cores and a memory included in a neural processor in accordance with some embodiments of the disclosure.
6 7 FIGS.and 16 FIG. 100 100 120 7 14 100 120 100 Referring to, the core array CoA may include a plurality of neural cores. Each of the plurality of neural coresmay include an L0 memory. Although FIGS.toshow that each of the neural coresincludes only the L0 memory, this is merely for easily describing the transmission paths of data, and embodiments are not limited thereto. The description of other components included in the neural corewill be described later with reference to.
1 2 In addition, the core array CoA may include a first data line D_Lthrough which data are transmitted in a first direction, and a second data line D_Lthrough which data are transmitted in a second direction opposite to the first direction. The first direction and the second direction may be referred to as a forward direction and a backward direction, respectively. In some embodiments, the core array CoA can transmit data in both directions.
100 1 100 2 100 1 2 The plurality of neural coresmay be connected in series with each other via the first data line D_L. Further, the plurality of neural coresmay be connected in series with each other via the second data line D_L. In other words, the plurality of neural coresmay have a structure in which they are connected in series with each other by the first data line D_Land the second data line D_L.
1 400 2 400 400 1 2 100 400 100 1 100 100 1 100 2 120 100 1 120 1 120 100 2 120 2 The first data line D_Lmay connect the memoryand the core array CoA. Further, the second data line D_Lmay connect the memoryand the core array CoA. The connection between the memoryand the core array CoA may be directly made via the first data line D_Land the second data line D_L, or indirectly made via a local interconnector to be described later. For the convenience of description, the neural coreconnected to the memoryis defined as a first neural core_, and the neural coreconnected in series with the first neural core_is defined as a second neural core_. In addition, the L0 memoryincluded in the first neural core_is defined as a first L0 memory_, and the L0 memoryincluded in the second neural core_is defined as a second L0 memory_. However, the selection of these terms is merely for the convenience of description, and the embodiments are not limited to these terms.
1 120 2 120 1 1 120 1 2 2 120 1 3 1 120 2 4 2 120 2 In some embodiments, the core array CoA may include a connection line connecting the first data line D_Land the L0 memoryand a connection line connecting the second data line D_Land the L0 memory. For example, the core array CoA may include a first connection line C_Lconnecting the first data line D_Land the first L0 memory_, and a second connection line C_Lconnecting the second data line D_Land the first L0 memory_. Furthermore, the core array CoA may include a third connection line C_Lconnecting the first data line D_Land the second L0 memory_, and a fourth connection line C_Lconnecting the second data line D_Land the second L0 memory_.
100 1 100 2 1 2 120 1 100 1 1 1 120 1 100 1 2 2 120 2 100 2 1 3 120 2 100 2 2 4 In some embodiments, the first neural core_and the second neural core_may be connected in series with each other via the first data line D_Land the second data line D_L. In addition, the first L0 memory_included in the first neural core_may be connected to the first data line D_Lvia the first connection line C_L. Further, the first L0 memory_included in the first neural core_may be connected to the second data line D_Lvia the second connection line C_L. Furthermore, the second L0 memory_included in the second neural core_may be connected to the first data line D_Lvia the third connection line C_L. Moreover, the second L0 memory_included in the second neural core_may be connected to the second data line D_Lvia the fourth connection line C_L.
300 300 1000 300 According to some embodiments, the core array CoA may include controllable ports Ctrlb_port that can be controlled on/off by software, firmware, or the task manager. According to some embodiments, the controllable ports Ctrlb_port may be implemented via hardware, software, or firmware. According to some embodiments, the task managerof the neural processormay configure data paths by controlling on/off of the controllable ports Ctrlb_port included in each core array CoA via a descriptor, but embodiments are not limited thereto. For example, in some other embodiments, data paths may also be configured by controlling the controllable ports Ctrlb_port via software or firmware. In the following, it will be described that the task managercontrols the controllable ports Ctrlb_port and configures data paths accordingly for convenience of description.
1 2 400 100 100 120 300 300 The controllable ports Ctrlb_port may be installed in the first data line D_L, the second data line D_L, and the connection line, and may set data movement paths. In some embodiments, the controllable ports Ctrlb_port may be disposed between the memoryand the neural core, between the plurality of neural cores, and between the connection line and the L0 memory, and the task managermay appropriately control the controllable ports Ctrlb_port to thereby configure data movement paths. In some embodiments, the task managermay fix the configured data paths or may reconfigure the data paths in real time.
1 400 100 1 100 1 100 2 300 400 100 1 400 100 1 300 100 1 100 2 100 1 100 2 For example, the first data line D_Lmay include the controllable ports Ctrlb_port between the memoryand the first neural core_and between the first neural core_and the second neural core_. That is, the task managermay configure a data movement path in a first direction between the memoryand the first neural core_by controlling the controllable port Ctrlb_port between the memoryand the first neural core_. Further, the task managermay configure a data movement path in the first direction between the first neural core_and the second neural core_by controlling the controllable ports Ctrlb_port between the first neural core_and the second neural core_.
2 400 100 1 100 1 100 2 300 400 100 1 400 100 1 300 100 1 100 2 100 1 100 2 In addition, for example, the second data line D_Lmay include the controllable ports Ctrlb_port between the memoryand the first neural core_and between the first neural core_and the second neural core_. That is, the task managermay configure a data movement path in a second direction between the memoryand the first neural core_by controlling the controllable port Ctrlb_port between the memoryand the first neural core_. Further, the task managermay configure a data movement path in the second direction between the first neural core_and the second neural core_by controlling the controllable ports Ctrlb_port between the first neural core_and the second neural core_.
1 1 120 1 300 1 120 1 1 120 1 2 3 4 300 120 160 120 120 16 FIG. Further, for example, the first connection line C_Lmay include a controllable port Ctrlb_port between the first data line D_Land the first 10 memory_. That is, the task managermay configure a data movement path in the first direction between the first data line D_Land the first L0 memory_by controlling the controllable port Ctrlb_port between the first data line D_Land the first L0 memory_. Similarly, each of the second connection line C_L, the third connection line C_L, and the fourth connection line C_Lmay include a controllable port Ctrlb_port, and the task managermay configure data movement paths in the first direction or the second direction by controlling the controllable ports Ctrlb_port. The data movement to the L0 memorymay mean, but not limited to, that computation is performed in a processing unit (in) corresponding to the L0 memoryfor the convenience of description. In some embodiments, there may also be cases in which even if data are provided to the L0 memory, it may be outputted without computation.
400 400 400 1 1 2 2 1 2 1 2 Hereinafter, embodiments will be described with i-th unit (i=1 . . . N). The 1st unit may be a memory, the i-th unit (i=2 . . . N−1) may be a (i−1)-th neural core, and the N-th unit may be the memoryor a memory different from the memory. The i-th unit (i=2 . . . N−1) may have a first port on the first data line D_Lin the backward direction, a second port on the first data line D_Lin the forward direction, a third port on the second data line D_Lin the backward direction, a fourth port on the second data line D_Lin the forward direction, a fifth port between an L0 memory of the i-th unit and the first data line D_L, and a sixth port between the L0 memory of the i-th unit and the second data line D_L. In some embodiments, a single port on the first data line D_Lor the second data line D_Lbetween the i-th unit and the (i+1) unit may function as a port installed in the i-th unit and a port installed in the (i+1) unit. In some embodiments, the single port may be installed in the i-th unit or the (i+1)-th unit, or be disposed between the i-th unit and the (i+1)-th unit. A port between the i-th unit and the (i+1)-th unit or a port in the forward direction of the i-th may be referred to as one or more ports between the i-th unit and the (i+1)-th unit regardless of where the one or more ports are installed. Similarly, a port between the (i−1)-th unit and the i-th unit or a port in the backward direction of the i-th may be referred to as one or more ports between the (i−1)-th unit and the i-th unit regardless of where the one or more ports are installed.
1 2 1 4 1 2 1 4 The first data line D_L, the second data line D_L, the first to fourth connection lines C_Lto C_L, and the controllable ports Ctrlb_port may be referred to respectively as a first data line circuit, a second data line circuit, first to fourth connection line circuits, and controllable port circuits. However, for the sake of convenience, the terms are respectively unified as a first data line, a second data line, first to fourth connection lines, and controllable ports. In addition, each of the first data line D_L, the second data line D_L, the first to fourth connection lines C_Lto C_L, and the controllable ports Ctrlb_port may be implemented as a circuit or circuitry.
100 1 2 120 1 2 300 400 100 300 In summary, the plurality of neural coresincluded in the core array CoA in accordance with some embodiments may be connected in series with each other by the first data line D_Land the second data line D_L. In addition, each of the L0 memoriesmay be connected to each other via the first data line D_Land the second data line D_L. Moreover, the core array CoA may include the controllable ports Ctrlb_port capable of controlling on/off. Therefore, the task managermay configure movement paths for the data provided from the memoryto the plurality of neural coresby controlling on/off of the controllable ports Ctrlb_port. In the following, movement paths of data configured by the task managerwill be described by way of example.
8 8 FIGS.A andB 1000 100 100 are example diagrams for illustrating a first data path in accordance with some embodiments of the disclosure. For convenience of description, the description is given assuming that the neural processorincludes four neural cores, and it is apparent that embodiments are not limited to the number of neural cores. Further, in the following, descriptions that are identical to or similar to those described above will be omitted or simplified for the convenience of description.
6 7 8 FIGS.,, andA 1000 400 1 400 2 100 1 100 2 100 3 100 4 400 1 400 2 400 400 400 400 Referring to, the neural processormay include a first memory_, a second memory_, and a core array CoA. The core array CoA may include a first neural core_, a second neural core_, a third neural core_, and a fourth neural core_. The first memory_and the second memory_may be the same memory or different memories. In some embodiments, the term ‘data path’ may be defined to refer to a path through which the data outputted from the memoryis inputted to the memory. In some embodiments, the ‘data path’ may refer to a path through which data moves from the memoryto the core array CoA and from the core array CoA to the memory. However, the definition of such a term is for convenience of description, and the embodiments are not limited to such a term.
100 1 120 1 100 2 120 2 100 3 120 3 100 4 120 4 The first neural core_may include a first L0 memory_. The second neural core_may include a second L0 memory_. The third neural core_may include a third L0 memory_. The fourth neural core_may include a fourth L0 memory_.
300 1 100 1 1 400 1 100 1 300 1 120 1 1 120 1 100 1 1 100 1 1 1 300 1 100 1 According to some embodiments, the task managermay provide first input data I_Dto the first neural core_by turning on the controllable port Ctrlb_port on the first data line D_Lbetween the first memory_and the first neural core_. In addition, the task managermay provide the first input data I_Dto the first L0 memory_by controlling the controllable port Ctrlb_port. The first input data I_Dprovided to the first L0 memory_may be computed by a processing unit of the first neural core_to generate first output data O_D. In some embodiments, the first neural core_may generate the first output data O_Dby using the first input data I_D. The task managermay control the controllable ports Ctrlb_port to provide the first output data O_Dto the outside of the first neural core_.
300 1 2 100 2 300 2 120 2 300 1 100 1 100 2 2 120 2 100 2 2 100 2 2 2 300 2 100 2 The task managermay control the controllable ports Ctrlb_port to provide the first output data O_Das second input data I_Dto the second neural core_. In addition, the task managermay control the controllable ports Ctrlb_port to provide the second input data I_Dto the second L0 memory_. In some embodiments, the task managermay turn on the controllable port Ctrlb_port on the first data line D_Lbetween the first neural core_and the second neural core_. Hereinafter, explanation on how to control the controllable ports Ctrlb_port will be omitted for convenience, since the person of ordinary skill in the art can induce how to control. The second input data I_Dprovided to the second L0 memory_may be computed by a processing unit of the second neural core_to generate second output data O_D. In some embodiments, the second neural core_may generate the second output data O_Dby using the second input data I_D. The task managermay control the controllable ports Ctrlb_port to provide the second output data O_Dto the outside of the second neural core_.
300 2 3 120 3 100 3 3 120 3 100 3 3 300 3 100 3 Likewise, the task managermay control the controllable ports Ctrlb_port to provide the second output data O_Das third input data I_Dto the third L0 memory_of the third neural core_. The third input data I_Dprovided to the third L0 memory_may be computed by a processing unit of the third neural core_to generate third output data O_D. The task managermay control the controllable ports Ctrlb_port to provide the third output data O_Dto the outside of the third neural core_.
300 3 4 120 4 100 4 4 120 4 100 4 4 300 4 100 4 In addition, the task managermay control the controllable ports Ctrlb_port to provide the third output data O_Das fourth input data I_Dto the fourth L0 memory_of the fourth neural core_. The fourth input data I_Dprovided to the fourth L0 memory_may be computed by a processing unit of the fourth neural core_to generate fourth output data O_D. The task managermay control the controllable ports Ctrlb_port to provide the fourth output data O_Dto the outside of the fourth neural core_.
300 300 400 1 120 1 120 1 120 2 120 2 120 3 120 3 120 4 That is, the task managermay move data in a first direction by controlling the controllable ports Ctrlb_port. For example, the task managermay move data from the first memory_to the first L0 memory_, from the first L0 memory_to the second L0 memory_, from the second L0 memory_to the third L0 memory_, and from the third L0 memory_to the fourth L0 memory_, by controlling the controllable ports Ctrlb_port.
6 7 8 FIGS.,, andB 300 4 5 100 3 300 5 120 3 5 120 3 100 3 5 100 3 5 5 300 5 100 3 Referring to, the task managermay control the controllable ports Ctrlb_port to provide the fourth output data O_Das fifth input data I_Dto the third neural core_. In addition, the task managermay control the controllable port Ctrlb_port to provide the fifth input data I_Dto the third L0 memory_. The fifth input data I_Dprovided to the third L0 memory_may be computed by the processing unit of the third neural core_to generate fifth output data O_D. In some embodiments, the third neural core_may generate the fifth output data O_Dby using the fifth input data I_D. The task managermay control the controllable ports Ctrlb_port to provide the fifth output data O_Dto the outside of the third neural core_.
300 5 6 120 2 100 2 6 120 2 100 2 6 300 6 100 2 Similarly, the task managermay control the controllable ports Ctrlb_port to provide the fifth output data O_Das sixth input data I_Dto the second L0 memory_of the second neural core_. The sixth input data I_Dprovided to the second L0 memory_may be computed by the processing unit of the second neural core_to generate sixth output data O_D. The task managermay control the controllable ports Ctrlb_port to provide the sixth output data O_Dto the outside of the second neural core_.
300 6 7 120 1 100 1 7 120 1 100 1 7 300 7 100 1 300 7 400 1 In addition, the task managermay control the controllable ports Ctrlb_port to provide the sixth output data O_Das seventh input data I_Dto the first L0 memory_of the first neural core_. The seventh input data I_Dprovided to the first L0 memory_may be computed by the processing unit of the first neural core_to generate seventh output data O_D. The task managermay control the controllable ports Ctrlb_port to provide the seventh output data O_Dto the outside of the first neural core_. The task managermay control to store the seventh output data O_Din the first memory_.
300 300 120 4 120 3 120 3 120 2 120 2 120 1 120 1 400 1 That is, the task managermay move data in a second direction by controlling the controllable ports Ctrlb_port. For example, the task managermay move data from the fourth L0 memory_to the third L0 memory_, from the third L0 memory_to the second L0 memory_, from the second L0 memory_to the first L0 memory_, and from the first L0 memory_to the first memory_, by controlling the controllable ports Ctrlb_port.
7 400 1 120 1 100 1 120 1 100 1 120 2 100 1 300 100 1 100 4 300 300 100 100 1000 100 According to some embodiments, the seventh output data O_Dprovided to the first memory_may be provided again as input data to the first L0 memory_of the first neural core_. The input data provided to the first L0 memory_may be computed by the processing unit of the first neural core_to generate output data. The generated output data may be provided to the second L0 memory_of the second neural core_. In some embodiments, the task managermay configure a first data path so that data repeatedly circulates through the first neural core_to the fourth neural core_. In some embodiments, the task managermay configure data paths in a theoretically infinite length by repeatedly configuring the first data path. Therefore, the task managermay repeatedly configure the first data path so that as many neural coresas needed for computation can be used regardless of the physical number of the actual neural cores. That is, the neural processorin accordance with some embodiments may configure data paths in an infinite length even with a small number of neural cores.
300 400 1 120 4 120 1 120 2 120 3 120 4 400 1 120 3 120 2 120 1 In summary, the task managermay configure the first data path by controlling the controllable ports Ctrlb_port included in the core array CoA. As described above, the first data path may include a path through which data move in the first direction from the first memory_to the fourth L0 memory_by way of the first L0 memory_, the second L0 memory_, and the third L0 memory_, and a path through which data move in the second direction from the fourth L0 memory_to the first memory_by way of the third L0 memory_, the second L0 memory_and the first L0 memory_.
9 FIG. is an example diagram for illustrating a second data path and a third data path in accordance with some embodiments of the disclosure.
6 7 9 FIGS.,, and 1000 400 1 400 2 100 1 100 2 100 3 100 4 400 1 400 2 Referring to, the neural processormay include a first memory_, a second memory_, and a core array CoA. The core array CoA may include a first neural core_, a second neural core_, a third neural core_, and a fourth neural core_, and each neural core may include an L0 memory. The first memory_and the second memory_may be the same memory or different memories.
300 400 1 120 1 120 1 120 2 120 2 120 3 120 3 120 4 120 4 400 2 According to some embodiments, the task managermay configure a second data path in the first direction by controlling the controllable ports Ctrlb_port, so that data are provided from the first memory_to the first L0 memory_, from the first L0 memory_to the second L0 memory_, from the second L0 memory_to the third L0 memory_, from the third L0 memory_to the fourth L0 memory_, and from the fourth L0 memory_to the second memory_.
300 400 1 400 2 100 1 100 4 120 100 In some embodiments, the task managermay provide data in the first direction from the first memory_toward the second memory_by controlling the controllable ports Ctrlb_port, so that data are computed by way of the first neural core_through the fourth neural core_. For example, it is assumed that the data provided to each L0 memoryare computed by a processing unit included in each neural core.
300 400 1 120 1 100 1 300 120 2 100 2 300 120 3 100 3 300 120 4 100 4 300 400 2 400 2 If described in terms of data computation, the task managermay control to provide first data outputted from the first memory_to the first L0 memory_. The first neural core_may generate second data by computing the first data. The task managermay control to provide the second data to the second L0 memory_. The second neural core_may generate third data by computing the second data. The task managermay control to provide the third data to the third L0 memory_. The third neural core_may generate fourth data by computing the third data. The task managermay control to provide the fourth data to the fourth L0 memory_. The fourth neural core_may generate fifth data by computing the fourth data. The task managermay provide the fifth data to the second memory_and control the second memory_to store the fifth data.
300 400 2 120 4 120 4 120 3 120 3 120 2 120 2 120 1 120 1 400 1 Further, the task managermay configure a third data path in the second direction by controlling the controllable ports Ctrlb_port, so that data are provided from the second memory_to the fourth L0 memory_, from the fourth L0 memory_to the third L0 memory_, from the third L0 memory_to the second L0 memory_, from the second L0 memory_to the first L0 memory_, and from the first L0 memory_to the first memory_.
300 400 2 400 1 100 4 100 1 In some embodiments, the task managermay provide data in the second direction from the second memory_toward the first memory_by controlling the controllable ports Ctrlb_port, so that data are computed by way of the fourth neural core_through the first neural core_.
300 400 2 120 4 100 4 300 120 3 100 3 300 120 2 100 2 300 120 1 100 1 300 400 1 400 1 If described in terms of data computation, the task managermay control to provide sixth data outputted from the second memory_to the fourth L0 memory_. The fourth neural core_may generate seventh data by computing the sixth data. The task managermay control to provide the seventh data to the third L0 memory_. The third neural core_may generate eighth data by computing the seventh data. The task managermay control to provide the eighth data to the second L0 memory_. The second neural core_may generate ninth data by computing the eighth data. The task managermay control to provide the ninth data to the first L0 memory_. The first neural core_may generate tenth data by computing the ninth data. The task managermay provide the tenth data to the first memory_and control the first memory_to store the tenth data.
300 400 1 400 2 120 1 120 2 120 3 120 4 400 2 400 1 120 4 120 3 120 2 120 1 In summary, the task managermay configure the second data path and the third data path by controlling the controllable ports Ctrlb_port included in the core array CoA. The second data path may be a path through which data move in the first direction from the first memory_to the second memory_by way of the first L0 memory_, the second L0 memory_, the third L0 memory_, and the fourth L0 memory_. In addition, the third data path may be a path through which data move in the second direction from the second memory_to the first memory_by way of the fourth L0 memory_, the third L0 memory_, the second L0 memory_, and the first L0 memory_.
10 FIG. is an example diagram for illustrating a fourth data path and a fifth data path in accordance with some embodiments of the disclosure.
6 7 10 FIGS.,, and 1000 400 1 400 2 1 2 1 2 1 100 1 100 2 100 3 2 100 4 400 1 400 2 Referring to, the neural processormay include a core array CoA, a first memory_, and a second memory_. The core array CoA may include a first computation group CoG_and a second computation group CoG_. The first computation group CoG_and the second computation group CoG_may execute different programs, applications, or computations. The first computation group CoG_may include a first neural core_, a second neural core_, and a third neural core_, and the second computation group CoG_may include a fourth neural core_. Each neural core may include an L0 memory. The first memory_and the second memory_may be the same memory or different memories.
300 1 300 400 1 120 1 120 1 120 2 120 2 120 3 120 3 120 2 120 2 120 1 120 1 400 1 According to some embodiments, the task managermay configure a fourth data path by controlling the controllable ports Ctrlb_port. The fourth data path may be a data path for the first computation group CoG_. The task managermay configure the fourth data path by configuring a data movement path in a first direction in which data are provided from the first memory_to the first L0 memory_, from the first L0 memory_to the second L0 memory_, and from the second L0 memory_to the third L0 memory_, and a data movement path in a second direction in which data are provided from the third L0 memory_to the second L0 memory_, from the second L0 memory_to the first L0 memory_, and from the first L0 memory_to the first memory_.
300 100 1 100 3 100 3 100 2 100 1 400 1 120 100 In some embodiments, the task managermay configure the fourth data path by controlling the controllable ports Ctrlb_port, so that data are computed by way of the first neural core_through the third neural core_, and the data computed in the third neural core_are computed again by way of the second neural core_and the first neural core_and stored in the first memory_. For example, it is assumed that the data provided to each L0 memoryare computed by a processing unit included in each neural core.
300 400 1 120 1 100 1 300 120 2 100 2 300 120 3 100 3 300 120 2 100 2 300 120 1 100 1 300 400 1 If described in terms of data computation, the task managermay control to provide first data outputted from the first memory_to the first L0 memory_. The first neural core_may generate second data by computing the first data. The task managermay control to provide the second data to the second L0 memory_. The second neural core_may generate third data by computing the second data. The task managermay control to provide the third data to the third L0 memory_. The third neural core_may generate fourth data by computing the third data. The task managermay control to provide the fourth data to the second L0 memory_. The second neural core_may generate fifth data by computing the fourth data. The task managermay control to provide the fifth data to the first L0 memory_. The first neural core_may generate sixth data by computing the fifth data. The task managermay control to store the sixth data in the first memory_.
300 2 300 400 2 120 4 120 4 400 2 Further, the task managermay configure a fifth data path by controlling the controllable ports Ctrlb_port. The fifth data path may be a data path for the second computation group CoG_. The task managermay configure the fifth data path by configuring a data movement path in the second direction in which data are provided from the second memory_to the fourth L0 memory_, and a data movement path in the first direction in which data are provided from the fourth L0 memory_to the second memory_again.
300 400 2 100 4 400 2 120 100 In some embodiments, the task managermay configure the fifth data path by controlling the controllable ports Ctrlb_port, so that the data provided from the second memory_are computed in the fourth neural core_and stored again in the second memory_. For example, it is assumed that the data provided to each L0 memoryare computed by a processing unit included in each neural core.
300 400 2 120 4 100 4 300 400 2 If described in terms of data computation, the task managermay control to provide seventh data outputted from the second memory_to the fourth L0 memory_. The fourth neural core_may generate eighth data by computing the seventh data. The task managermay control to store the eighth data in the second memory_.
1 2 300 1 2 400 1 120 1 120 2 120 3 120 3 120 2 120 1 400 1 400 2 120 4 120 4 400 2 In summary, the first computation group CoG_and the second computation group CoG_included in the core array CoA may have different data paths. That is, the task managermay configure the fourth data path of the first computation group CoG_and the fifth data path of the second computation group CoG_by controlling the controllable ports Ctrlb_port included in the core array CoA. The fourth data path may include a path through which data are moved in the first direction from the first memory_through the first L0 memory_and the second L0 memory_, to the third L0 memory_, and a path through which data are moved in the second direction from the third L0 memory_through the second L0 memory_and the first L0 memory_, to the first memory_. In addition, the fifth data path may include a path through which data are moved in the second direction from the second memory_to the fourth L0 memory_, and a path through which data are moved in the first direction from the fourth L0 memory_to the second memory_.
11 FIG. is an example diagram for illustrating a sixth data path and a seventh data path in accordance with some embodiments of the disclosure.
6 7 11 FIGS.,, and 1000 400 1 400 2 100 1 100 2 100 3 100 4 400 1 400 2 Referring to, the neural processormay include a first memory_, a second memory_, and a core array CoA. The core array CoA may include a first neural core_, a second neural core_, a third neural core_, and a fourth neural core_, and each neural core may include an L0 memory. The first memory_and the second memory_may be the same memory or different memories.
300 400 1 120 1 120 1 120 2 120 2 400 2 According to some embodiments, the task managermay configure a sixth data path in the first direction by controlling the controllable ports Ctrlb_port, so that data are provided from the first memory_to the first L0 memory_, from the first L0 memory_to the second L0 memory_, and from the second L0 memory_to the second memory_.
300 400 1 400 2 100 1 100 2 120 100 1 100 2 In some embodiments, the task managermay control that data are provided in the first direction from the first memory_toward the second memory_but the data are computed only in the first neural core_and the second neural core_by controlling the controllable ports Ctrlb_port. For example, it is assumed that the data provided to each L0 memoryare computed by a processing unit included in the first neural core_and the second neural core_.
300 400 1 120 1 100 1 300 120 2 100 2 300 400 2 400 2 If described in terms of data computation, the task managermay control to provide first data outputted from the first memory_to the first L0 memory_. The first neural core_may generate second data by computing the first data. The task managermay control to provide the second data to the second L0 memory_. The second neural core_may generate third data by computing the second data. The task managermay provide the third data to the second memory_and control the second memory_to store the third data.
1 100 1 100 2 100 1 100 2 1 100 3 100 4 120 2 400 2 300 1 1 300 1000 100 3 100 4 According to some embodiments, a part of a first data line D_Lpassing through the first neural core_and the second neural core_may be used as a data computation path that allows computation of data to be performed in the first neural core_and the second neural core_, and the other part of the first data line D_Lpassing through the third neural core_and the fourth neural core_may be used as a data bus through which data pass from the second L0 memory_to the second memory_. In some embodiments, the task managerconfigures the sixth data path and may use part of the first data line D_Las the data computation path and the other part of the first data line D_Las the data bus. In this case, the task managercan also minimize the power consumption of the neural processorby adjusting the power so that the third neural core_and the fourth neural core_, in which computation is not performed, are not driven.
300 400 2 120 2 120 2 120 1 120 1 400 1 Furthermore, the task managermay configure the seventh data path in the second direction by controlling the controllable ports Ctrlb_port, so that data are provided from the second memory_to the second L0 memory_, from the second L0 memory_to the first L0 memory_, and from the first L0 memory_to the first memory_.
300 400 2 400 1 100 2 100 1 120 100 1 100 2 In some embodiments, the task managermay control that data are provided in the second direction from the second memory_toward the first memory_but the data are computed only in the second neural core_and the first neural core_by controlling the controllable ports Ctrlb_port. For example, it is assumed that the data provided to each L0 memoryare computed by a processing unit included in the first neural core_and the second neural core_.
300 400 2 120 2 100 2 300 120 1 100 1 300 400 1 400 1 If described in terms of data computation, the task managermay control to provide fourth data outputted from the second memory_to the second L0 memory_. The second neural core_may generate fifth data by computing the fourth data. The task managermay control to provide the fifth data to the first L0 memory_. The first neural core_may generate sixth data by computing the fifth data. The task managermay provide the sixth data to the first memory_and control the first memory_to store the sixth data.
2 100 4 100 3 400 2 120 2 2 100 2 100 1 100 2 100 1 300 2 2 300 1000 100 3 100 4 According to some embodiments, a part of a second data line D_Lpassing through the fourth neural core_and the third neural core_may be used as a data bus through which data pass from the second memory_to the second L0 memory_, and the other part of the second data line D_Lpassing through the second neural core_and the first neural core_may be used as a data computation path that allows the computation of data to be performed in the second neural core_and the first neural core_. In some embodiments, the task managerconfigures the seventh data path, and may use part of the second data line D_Las the data computation path and the other part of the second data line D_Las the data bus. In this case, the task managercan also minimize the power consumption of the neural processorby adjusting the power so that the third neural core_and the fourth neural core_, in which computation is not performed, are not driven.
300 1000 1 2 100 300 100 100 100 1 100 2 300 120 1 120 2 120 3 120 4 According to some embodiments, the task managercan enhance the security of the neural processorby using only parts of the first data line D_Land the second data line D_Las the data computation path. Sensitive information such as personal information should be computed and handled only in particular neural cores. In this case, the task managermay control data to be computed only in particular neural coresincluded in the core array CoA, and that the rest of the neural corescannot check or access the data, by controlling the controllable ports Ctrlb_port. For example, if particular data needs to be handled by the first neural core_and the second neural core_only, the task managermay control that the data are provided only to the first L0 memory_and the second L0 memory_, and may control that the third L0 memory_and the fourth L0 memory_cannot check the data, by controlling the controllable ports Ctrlb_port.
300 400 1 120 1 120 2 400 2 400 2 120 2 120 1 400 1 In summary, the task managermay configure the sixth data path and the seventh data path by controlling the controllable ports Ctrlb_port included in the core array CoA. The sixth data path may be a path through which data are moved in the first direction from the first memory_through the first L0 memory_and the second L0 memory_, to the second memory_. In addition, the seventh data path may be a path through which data are moved in the second direction from the second memory_through the second L0 memory_and the first L0 memory_, to the first memory_.
300 7 11 FIGS.to Some examples of the data paths configured by the task managerhave been described referring to. However, the embodiments of the disclosure are not limited to these examples, and those having ordinary skill in the art of the disclosure may be able to configure data paths that have not been described separately herein by appropriately modifying the data paths without departing from the scope of the disclosure.
100 100 100 100 100 1 2 According to some embodiments, the neural coresincluded in the core array CoA are connected in series. If it is necessary to add neural coresto the core array CoA, simple expansion is possible without the need for separate design changes or addition of lines because all of the neural coresincluded in the core array CoA are connected in series. For example, since new neural corescan be added to the core array CoA by simply connecting the neural coresto be newly added with the first data line D_Land the second data line D_Lin series and connecting the controllable ports Ctrlb_port, there is an advantage of being highly scalable.
300 300 300 1000 According to some embodiments, the task managermay configure various data paths by using the controllable ports Ctrlb_port. In some embodiments, the task managercan configure data paths, and thus can design and reflect appropriate data paths according to a flow of data. That is, the task managercan configure the most optimized data paths according to the type of computations, and thus, unnecessary data movement can be reduced, thereby increasing the computation efficiency of the neural processor.
12 FIG. is a diagram for illustrating data lines and connection lines connecting neural cores and memories included in a neural processor in accordance with some embodiments of the disclosure. For the convenience of description, descriptions that are identical to or similar to those described above will be omitted or simplified.
6 12 FIGS.and 1 2 Referring to, the core array CoA may include a plurality of first data lines D_Lthrough which data are transmitted in a first direction, and a plurality of second data lines D_Lthrough which data are transmitted in a second direction. In some embodiments, the core array CoA can transmit data in both directions.
100 1 2 100 1 100 2 100 1 2 The core array CoA may include a plurality of neural coresand, the first data lines D_L, and the second data lines D_L. The plurality of neural coresmay be connected in series with each other via the first data lines D_L. Further, the plurality of neural coresmay be connected in series with each other via the second data lines D_L. In some embodiments, the plurality of neural coresmay have a structure in which they are connected in series with each other by the first data lines D_Land the second data lines D_L.
1 2 1 2 The first data line D_Land the second data line D_Lmay each include a plurality of data lines. In some embodiments, the first data lines D_Lmay include a plurality of data lines through which data are transmitted in the first direction. Also, the second data lines D_Lmay include a plurality of data lines through which data are transmitted in the second direction.
1 120 2 120 1 2 1 120 2 120 In some embodiments, the core array CoA may include a plurality of connection lines connecting the first data line D_Land the L0 memoryand a plurality of connection lines connecting the second data line D_Land the L0 memory. As described above, since each of the first data lines D_Land the second data lines D_Lincludes a plurality of data lines, the connection lines connecting the first data line D_Land the L0 memorymay also be configured in plurality, and the connection lines connecting the second data line D_Land the L0 memorymay also be configured in plurality.
1 2 According to some embodiments, the core array CoA may include controllable ports Ctrlb_port that can be controlled on/off by software or firmware. The controllable ports Ctrlb_port may be installed on the first data lines D_L, the second data lines D_L, and the connection lines, and may set data movement paths.
300 1 300 1 300 1 According to some embodiments, the task managermay turn on/off at least some of the plurality of data lines included in the first data lines D_Lby controlling the controllable ports Ctrlb_port. For example, the task managermay turn off some of the plurality of data lines included in the first data lines D_Lby controlling the controllable ports Ctrlb_port. For another example, the task managermay turn on all of the plurality of data lines included in the first data lines D_Lby controlling the controllable ports Ctrlb_port.
300 2 300 2 300 2 Likewise, the task managermay turn on/off at least some of the plurality of data lines included in the second data lines D_Lby controlling the controllable ports Ctrlb_port. For example, the task managermay turn off some of the plurality of data lines included in the second data lines D_Lby controlling the controllable ports Ctrlb_port. For another example, the task managermay turn on all of the plurality of data lines included in the second data lines D_Lby controlling the controllable ports Ctrlb_port.
300 1 2 300 1 2 300 1 2 300 13 13 FIGS.A andB According to some embodiments, the task managercan prevent unnecessary power consumption by controlling at least some of the plurality of data lines included in the first data lines D_Land the second data lines D_Laccording to bandwidths of the data. For example, if a relatively high bandwidth is required for data transmission, the task managermay control the controllable ports Ctrlb_port to turn on all of the plurality of data lines included in the first data lines D_Land the second data lines D_Land use all of the plurality of data lines for data transmission. If a relatively low bandwidth is required for data transmission, the task managermay control the controllable ports Ctrlb_port to turn on some of the plurality of data lines included in the first data line D_Land the second data line D_Land turn off the rest of the data lines, thereby reducing unnecessary power consumption. In some embodiments, the task managercan minimize waste of power consumption by turning off unnecessary data transmission lines according to the bandwidths required for data transmission. Further reference is made tofor an illustrative description.
13 13 FIGS.A andB 13 FIG.A 13 FIG.B are example diagrams for illustrating on/off of data lines in accordance with some embodiments of the disclosure.describes a case where the bandwidth required for data transmission is relatively low, anddescribes a case where the bandwidth required for data transmission is relatively high.
6 12 13 FIGS.,, andA 1000 400 1 400 2 100 1 100 2 100 3 100 4 100 1 120 1 100 2 120 2 100 3 120 3 100 4 120 4 400 1 400 2 Referring to, the neural processormay include a first memory_, a second memory_, and a core array CoA. The core array CoA may include a first neural core_, a second neural core_, a third neural core_, and a fourth neural core_, and each neural core may include an L0 memory. In some embodiments, the first neural core_may include a first L0 memory_, the second neural core_may include a second L0 memory_, the third neural core_may include a third L0 memory_, and the fourth neural core_may include a fourth L0 memory_. The first memory_and the second memory_may be the same memory or different memories.
1 2 1 11 12 2 21 22 In addition, the core array CoA may include first data lines D_Land second data lines D_L. The first data lines D_Lmay include a (1-1)th data line D_Land a (1-2)th data line D_L. Further, the second data line D_Lmay include a (2-1)th data line D_Land a (2-2)th data line D_L.
11 12 21 22 11 12 21 22 The (1-1)th data line D_L, the (1-2)th data line D_L, the (2-1)th data line D_L, and the (2-2)th data line D_Lmay be referred to respectively as a (1-1)th data line circuit, a (1-2)th data line circuit, a (2-1)th data line circuit, and a (2-2)th data line circuit. However, for the sake of convenience, the terms are respectively unified as a (1-1)th data line, a (1-2)th data line, a (2-1)th data line, and a (2-2)th data line. In addition, each of the (1-1)th data line D_L, the (1-2)th data line D_L, the (2-1)th data line D_L, and the (2-2)th data line D_Lmay be implemented as a circuit or circuitry.
300 1 2 300 12 1 22 2 300 11 21 300 12 22 According to some embodiments, the task managermay turn off some of the data lines included in the first data lines D_Land the second data lines D_Lby controlling the controllable ports Ctrlb_port. For example, the task managermay turn off the (1-2)th data line D_Lincluded in the first data lines D_Land the (2-2)th data line D_Lincluded in the second data lines D_Lby controlling the controllable ports Ctrlb_port. In some embodiments, the task managermay configure data paths to enable data transmission by using only the (1-1)th data line D_Land the (2-1)th data line D_Lby controlling the controllable ports Ctrlb_port. That is, the task managercan reduce unnecessary power consumption by turning off the (1-2)th data line D_Land the (2-2)th data line D_Las needed.
13 FIG.B 300 1 2 300 11 12 1 21 22 2 300 11 12 21 22 300 11 12 21 22 Referring further to, the task managermay turn on all the data lines included in the first data lines D_Land the second data lines D_Lby controlling the controllable ports Ctrlb_port. For example, the task managermay turn on the (1-1)th data line D_Land the (1-2)th data line D_Lincluded in the first data lines D_L, and the (2-1)th data line D_Land the (2-2)th data line D_Lincluded in the second data lines D_Lby controlling the controllable ports Ctrlb_port. In some embodiments, the task managermay configure data paths to enable data transmission by using all of the (1-1)th data line D_L, the (1-2)th data line D_L, the (2-1)th data line D_L, and the (2-2)th data line D_Lby controlling the controllable ports Ctrlb_port. That is, the task managercan maximize transmission performance by turning on all of the (1-1)th data line D_L, the (1-2)th data line D_L, the (2-1)th data line D_L, and the (2-2)th data line D_Las needed.
14 FIG. is a diagram for illustrating data lines and connection lines connecting neural cores and memories included in a neural processor in accordance with some embodiments of the disclosure. For the convenience of description, descriptions that are identical to or similar to those described above will be omitted or simplified.
6 14 FIGS.and 1 2 1 1 1 2 1 Referring to, the core array CoA may include a first core array CoA_and a second core array CoA_. The first core array CoA_may include a first data line D_Lthrough which data are transmitted in a first direction. Further, the first core array CoA_may include a second data line D_Lthrough which data are transmitted in a second direction. In some embodiments, the first core array CoA_may be able to transmit data in both directions.
2 3 2 4 2 3 4 3 4 Also, the second core array CoA_may include a third data line D_Lthrough which data are transmitted in the first direction. In addition, the second core array CoA_may include a fourth data line D_Lthrough which data are transmitted in the second direction. In some embodiments, the second core array CoA_may be able to transmit data in both directions. In this case, the third data line D_Land the fourth data line D_Lmay be referred to respectively as a third data line circuit and a fourth data line circuit. However, for the sake of convenience, the terms are respectively unified as a third data line and a fourth data line. In addition, each of the third data line D_Land the fourth data line D_Lmay be implemented as a circuit or circuitry.
1 2 7 13 FIGS.toB The first core array CoA_and the second core array CoA_may be the core arrays CoA described with reference to.
400 500 1 500 2 500 400 1 2 500 1 400 2 500 2 400 1 500 The memorymay be connected to a local interconnection. In addition, the first core array CoA_may be connected to the local interconnection. Moreover, the second core array CoA_may be connected to the local interconnection. In some embodiments, data outputted from the memorymay be provided to the first core array CoA_and/or the second core array CoA_via the local interconnection. Further, data outputted from the first core array CoA_may be provided to the memoryand/or the second core array CoA_via the local interconnection. Moreover, data outputted from the second core array CoA_may be provided to the memoryand/or the first core array CoA_via the local interconnection.
500 200 300 500 400 200 300 500 6000 6000 3 FIG. The local interconnectionmay connect at least one core array CoA, the core global, and the task managerto one another. The local interconnectionmay be a path through which data move between the at least one core array CoA, the memory, the core global, and the task manager. The local interconnectionmay be connected to the global interconnectionofand transmit data to the global interconnection.
1000 400 400 500 500 That is, the neural processormay include the memoryand the plurality of core arrays CoA, and data movement may occur between the memoryand the plurality of core arrays CoA via the local interconnection. In addition, data movement between the plurality of core arrays CoA may also be performed via the local interconnection.
1 4 1 300 500 500 300 500 According to some embodiments, each of the first data line D_Lthrough the fourth data line D_Lmay include a plurality of data lines. For example, a description will be provided assuming a case in which the first data line D_Lincludes a (1-1)th data line and a (1-2)th data line. According to some embodiments, the task managermay control on/off of the (1-1)th data line and the (1-2)th data line according to bandwidths for transmitting data. If the bandwidth of the local interconnectionis greater than the bandwidth of the (1-1)th data line, a latency may increase due to a bottleneck when data are provided from the local interconnectionto the (1-1)th data line. Therefore, in this case, the task managermay turn on both the (1-1)th data line and the (1-2)th data line by controlling the controllable ports Ctrlb_port. If both the (1-1)th data line and the (1-2)th data line are turned on, the bottleneck occurring in the local interconnectioncan be minimized, and thus latency can be reduced accordingly.
500 500 300 1000 On the other hand, if the bandwidth of the local interconnectionis smaller than the bandwidth of the (1-1)th data line, no bottleneck may occur even if data are provided from the local interconnectionto the (1-1)th data line. In this case, the task managermay turn on the (1-1)th data line and turn off the (1-2)th data line by controlling the controllable ports Ctrlb_port. Through this, power consumption in the neural processorcan be minimized without affecting latency, and efficiency can thus be maximized.
300 1 2 500 1000 In some embodiments, the task managermay control at least some of the plurality of data lines included in the first data line D_Land the second data line D_Laccording to the bandwidth of the local interconnection. Through this, the efficiency of the neural processorin terms of power and latency can be maximized.
15 FIG. is a diagram for illustrating a hierarchical structure of a neural processing device in accordance with some embodiments of the disclosure.
15 FIG. 10 1000 1000 6000 Referring to, the neural core SoCmay include at least one neural processor. The neural processorsmay transmit data to each other via the global interconnection.
1000 100 100 100 100 Each of neural processorsmay include at least one neural core. The neural coremay be a unit of processing optimized for deep learning computation tasks. The neural coremay be a unit of processing corresponding to one operation of deep learning computation tasks. In some embodiments, a deep learning computation task can be represented by a sequential or parallel combination of multiple operations. Each of neural coresmay be a unit of processing capable of processing one operation, and may be a minimum computation unit that can be considered for scheduling from the viewpoint of a compiler.
The neural processing device in accordance with the embodiment may configure scales of the minimum computation unit considered from the viewpoint of compiler scheduling and the hardware unit of processing to be the same, so that fast and efficient scheduling and computation tasks can be performed.
That is, if a unit of processing into which hardware can be divided is too large compared to computation tasks, inefficiency of the computation tasks may occur in driving the unit of processing. Conversely, it is not appropriate to schedule a unit of processing that is a unit smaller than an operation, which is the minimum scheduling unit of the compiler, every time since a scheduling inefficiency may occur and hardware design costs may increase.
Therefore, by adjusting the scales of the scheduling unit of the compiler and the hardware unit of processing to be similar in the embodiment, it is possible to simultaneously satisfy the fast scheduling of computation tasks and the efficient execution of the computation tasks without wasting hardware resources.
16 FIG. 6 FIG. is a block diagram for illustrating the neural core ofin detail.
16 FIG. 100 110 120 130 140 150 160 Referring to, the neural coremay include a load/store unit (LSU), an L0 memory, a weight buffer, an activation LSU, an activation buffer, and a processing unit.
110 500 110 120 110 500 110 110 The LSUmay receive at least one of data, control signals, or synchronization signals from the outside via the local interconnectionand the L1 sync path. The LSUmay transmit at least one of the data, the control signals, or the synchronization signals received to the L0 memory. Similarly, the LSUmay transfer at least one of the data, the control signals, or the synchronization signals to the outside via the local interconnectionand the L1 sync path. In this case, the LSUmay be referred to as an LSU circuit, but for the sake of convenience, the terms are unified as an LSU. In addition, the LSUmay be implemented as a circuit or circuitry.
17 FIG. 16 FIG. is a block diagram for illustrating the LSU ofin detail.
17 FIG. 110 111 111 112 112 113 113 114 a b a b a b Referring to, the LSUmay include a local memory load unit (LMLU), a local memory store unit (LMSU), a neural core load unit (NCLU), a neural core store unit (NCSU), a load buffer LB, a store buffer SB, a load (LD) engine, a store (ST) engine, and a translation lookaside buffer (TLB).
111 111 112 112 113 113 111 111 112 112 113 113 a b a b a b a b a b a b The local memory load unit, the local memory store unit, the neural core load unit, the neural core store unit, the load engine, and the store enginemay be referred to respectively as a local memory load circuit, a local memory store circuit, a neural core load circuit, a neural core store circuit, a load engine circuit, and a store engine circuit. However, for the sake of convenience, the terms are respectively unified as a local memory load unit, a local memory store unit, a neural core load unit, a neural core store unit, a load engine, and a store engine. In addition, each of the local memory load unit, the local memory store unit, the neural core load unit, the neural core store unit, the load engine, and the store enginemay be implemented as a circuit or circuitry.
111 120 111 113 a a a The local memory load unitmay fetch a load instruction for the L0 memoryand issue the load instruction. When the local memory load unitprovides the issued load instruction to the load buffer LB, the load buffer LB may sequentially transmit memory access requests to the load engineaccording to the inputted order.
111 120 111 113 b b b Further, the local memory store unitmay fetch a store instruction for the L0 memoryand issue the store instruction. When the local memory store unitprovides the issued store instruction to the store buffer SB, the store buffer SB may sequentially transmit memory access requests to the store engineaccording to the inputted order.
112 100 112 113 a a a The neural core load unitmay fetch a load instruction for the neural coreand issue the load instruction. When the neural core load unitprovides the issued load instruction to the load buffer LB, the load buffer LB may sequentially transmit memory access requests to the load engineaccording to the inputted order.
112 100 112 113 b b b In addition, the neural core store unitmay fetch a store instruction for the neural coreand issue the store instruction. When the neural core store unitprovides the issued store instruction to the store buffer SB, the store buffer SB may sequentially transmit memory access requests to the store engineaccording to the inputted order.
113 500 113 114 113 114 a a a The load enginemay receive the memory access request and retrieve data via the local interconnection. In some embodiments, the load enginemay quickly find the data by using a translation table of a logical address and a physical address that has been used recently in the translation lookaside buffer. If the logical address of the load engineis not in the translation lookaside buffer, the address translation information may be found in another memory.
113 500 113 114 113 114 b b b The store enginemay receive the memory access request and retrieve data via the local interconnection. In some embodiments, the store enginemay quickly find the data by using a translation table of a logical address and a physical address that has been used recently in the translation lookaside buffer. If the logical address of the store engineis not in the translation lookaside buffer, the address translation information may be found in another memory.
113 113 a b The load engineand the store enginemay send synchronization signals to the L1 sync path. In some embodiments, the synchronization signal may indicate that the task has been completed.
16 FIG. 120 100 100 120 100 Referring toagain, the L0 memoryis a memory located inside the neural core, and may receive all input data required for the tasks by the neural corefrom the outside and store the input data temporarily. In addition, the L0 memorymay temporarily store the output data calculated by the neural corefor transmission to the outside.
120 150 140 120 160 140 120 163 164 120 120 The L0 memorymay transmit an input activation Act_In to the activation bufferand receive an output activation Act_Out via the activation LSU. The L0 memorymay directly transmit and receive data to and from the processing unit, in addition to the activation LSU. In some embodiments, the L0 memorymay exchange data with each of a processing element (PE) arrayand a vector unit. The L0 memorymay be a memory corresponding to the level of the neural core. In this case, the L0 memorymay be a private memory of the neural core that is not shared.
120 120 The L0 memorymay be a memory corresponding to the level of the neural core. In this case, the L0 memorymay be a private memory of the neural core.
120 120 120 110 130 140 160 The L0 memorymay transmit data such as activations or weights via a data path. The L0 memorymay exchange synchronization signals via an L0 sync path, which is a separate dedicated path. The L0 memorymay exchange synchronization signals with, for example, the LSU, the weight buffer, the activation LSU, and the processing unitvia the L0 sync path.
130 120 130 160 130 The weight buffermay receive a weight from the L0 memory. The weight buffermay transfer the weight to the processing unit. The weight buffermay temporarily store the weight before transferring the weight.
The input activation Act_In and the output activation Act_Out may refer to input values and output values of the layers of a neural network. In this case, if there are a plurality of layers in the neural network, the output value of the previous layer becomes the input value of the next layer, and thus, the output activation Act_Out of the previous layer may be utilized as the input activation Act_In of the next layer.
The weight may refer to a parameter that is multiplied by the input activation Act_In inputted in each layer. The weight is adjusted and confirmed in the deep learning training phase, and may be used to derive the output activation Act_Out via a fixed value in the inference phase.
140 120 150 150 140 The activation LSUmay transfer the input activation Act_In from the L0 memoryto the activation buffer, and the output activation Act_Out from the activation bufferto the on-chip buffer. In some embodiments, the activation LSUmay perform both load tasks and store tasks of the activation.
150 160 160 150 The activation buffermay provide the input activation Act_In to the processing unitand receive the output activation Act_Out from the processing unit. The activation buffermay temporarily store the input activation Act_In and the output activation Act_Out.
150 160 163 100 The activation buffermay quickly provide the activation to the processing unit, in particular, the PE array, which has a large quantity of calculations, and may quickly receive the activation, thereby increasing the calculation speed of the neural core.
160 160 160 The processing unitmay be a module that performs calculations. The processing unitmay perform not only one-dimensional calculations but also two-dimensional matrix calculations, i.e., convolution operations. The processing unitmay receive an input activation Act_In, multiply the input activation Act_In by a weight, and then add it to generate an output activation Act_Out.
18 FIG. 16 FIG. is a block diagram for illustrating the processing unit ofin detail.
16 FIG. 18 FIG. 160 163 164 161 162 Referring toand, the processing unitmay include a PE array, a vector unit, a column register, and a row register.
163 163 163 The PE arraymay receive the input activation Act_In and the weight and perform multiplication on them. In this case, each of the input activation Act_In and the weight may be in the form of matrices and calculated via convolution. Through this, the PE arraymay generate an output activation Act_Out. However, the embodiment is not limited thereto. The PE arraymay generate any types of outputs other than the output activation Act_Out as well.
163 163 1 163 1 163 1 The PE arraymay include at least one processing element (PE)_. The processing elements_may be aligned with each other so that each of the processing elements_may perform multiplication on one input activation Act_In and one weight.
163 163 The PE arraymay sum values for each multiplication to generate a subtotal. This subtotal may be utilized as an output activation Act_Out. The PE arrayperforms two-dimensional matrix multiplication, and thus, may be referred to as a 2D matrix compute unit.
164 164 163 160 100 The vector unitmay mainly perform one-dimensional calculations. The vector unit, together with the PE array, may perform deep learning calculations. Through this, the processing unitmay be specialized for necessary calculations. In some embodiments, each of the at least one neural corehas calculation modules that perform a large amount of two-dimensional matrix multiplications and one-dimensional calculations, and thus, can efficiently perform deep learning tasks.
161 1 161 1 163 1 The column registermay receive a first input I. The column registermay receive the first input I, and distribute them to each column of the processing elements_.
162 2 162 2 163 1 The row registermay receive a second input I. The row registermay receive the second input I, and distribute them to each row of the processing elements_.
1 2 1 1 2 The first input Imay be an input activation Act_In or a weight. The second input Imay be a value other than the first input Ibetween the input activation Act_In or the weight. Alternatively, the first input Iand the second input Imay be values other than the input activation Act_In and the weight.
19 FIG. 16 FIG. is a block diagram for illustrating the L0 memory ofin detail.
19 FIG. 120 121 122 Referring to, the L0 memorymay include a schedulerand one or more local memory banks.
120 121 113 122 122 a When data are stored in the L0 memory, the schedulermay receive data from the load engine. In this case, the local memory bankmay be allocated for the data in a round-robin manner. Accordingly, data may be stored in any one of the local memory banks.
120 121 122 113 113 500 121 121 b b In contrast to this, when data are loaded from the L0 memory, the schedulermay receive the data from the local memory bankand transmit the data to the store engine. The store enginemay store the data in the outside through the local interconnection. In this case, the schedulermay be referred to as a scheduler circuit, but for the sake of convenience, the terms are unified as a scheduler. In addition, the schedulermay be implemented as a circuit or circuitry.
20 FIG. 19 FIG. is a block diagram for illustrating the local memory bank ofin detail.
20 FIG. 122 122 1 122 2 Referring to, the local memory bankmay include a local memory bank controller_and a local memory bank cell array_.
122 1 122 122 1 The local memory bank controller_may manage read and write operations via the addresses of data stored in the local memory bank. In some embodiments, the local memory bank controller_may manage the input/output of data as a whole.
1222 122 2 122 1 The local memory bank cell arraymay be of a structure in which cells in which data are directly stored are arranged in rows and columns. The local memory bank cell array_may be controlled by the local memory bank controller_.
21 FIG. 1 FIG. 22 FIG. 21 FIG. is a block diagram for illustrating a flow of data and control signals of the neural processing device of, andis a block diagram for illustrating relationship between the command processor and the task managers of.
21 22 FIGS.and 1000 100 1000 300 700 300 7000 Referring to, the neural processormay include at least one neural core. Each neural processormay include a task managerand an L1 LSUtherein, respectively. The task managersmay exchange control signals and responses to the control signals with a command processorvia a control interconnection CI.
700 500 6100 400 2000 32 In contrast, the L1 LSUmay exchange data via a data interconnection and memory DIM. The data interconnection and memory DIM may include an interconnection for transmitting data and a memory in which data are shared. Specifically, the data interconnection and memory DIM may include a local interconnectionand a data channel. In addition, the data interconnection and memory DIM may include an L1 shared memory, a shared memory, and a volatile memory. However, the embodiment is not limited thereto.
300 7000 7000 300 300 7000 300 1000 1000 300 300 7000 The task managersmay be controlled by the command processor. That is, the command processormay transfer tasks to the task managersvia control signals, and the task managersmay transfer task completion reports to the command processor. At least one task managermay be included in the neural processor. Moreover, if the neural processorsare plural, the number of task managersmay get larger. Such a plurality of task managersmay all be controlled by the command processor.
23 FIG. is a block diagram for illustrating in detail the structure of the neural processing device in accordance with some embodiments of the disclosure.
23 FIG. 101 100 101 1111 111 2 111 3 111 4 1113 Referring to, a neural coremay have a CGRA structure, unlike a neural core. The neural coremay include an instruction memory, a CGRA L0 memory_, a PE array_, and a load/store unit (LSU)_. The PE arraymay include a plurality of processing elements interconnected by a mesh style network. The mesh style network may be two-dimensional, three-dimensional, or higher-dimensional. In the CGRA, the plurality of processing elements may be reconfigurable or programmable. The interconnection between the plurality of processing elements may be reconfigurable or programmable. In some embodiments, the interconnection between the plurality of processing elements may be statically reconfigurable or programmable when the interconnection is fixed after the plurality of processing elements are configurated or programed. In some embodiments, the interconnection between the plurality of processing elements may be dynamically reconfigurable or programmable when the interconnection is reconfigurable or programmable even after the plurality of processing elements are configurated or programed.
111 1 1111 111 3 111 3 1113 a The instruction memory_may receive and store instructions. The instruction memorymay sequentially store instructions internally, and provide the stored instructions to the PE array_. In this case, the instructions may instruct the operation of first type of a plurality of processing elements_included in each PE array.
1112 101 101 1112 101 1112 101 The CGRA L0 memorymay be located inside the neural core, receive all input data required for tasks of the neural core, and temporarily store the data. In addition, the CGRA L0 memorymay temporarily store output data calculated by the neural coreto transmit the data to the outside. The CGRA L0 memorymay serve as a cache memory of the neural core.
1112 111 3 111 2 101 111 2 111 3 The CGRA L0 memorymay send and receive data to and from the PE array_. The CGRA L0 memory_may be a memory corresponding to L0 (level 0) that is lower than L1. In this case, the L0 memory may be a private memory of the neural corethat is not shared. The CGRA L0 memory_may transmit data such as activations or weights, programs, and the like to the PE array_.
1113 111 3 1113 111 3 111 3 a b The PE arraymay be a module that performs calculations. The PE array_may perform not only one-dimensional calculations but also two-dimensional or higher matrix/tensor calculations. The PE arraymay include the first type of the plurality of processing elements_and a second type of a plurality of processing elements_therein.
111 3 111 3 111 3 111 3 111 3 111 3 111 3 111 3 a b a b a b a b The first type of the plurality of processing elements_and the second type of the plurality of processing elements_may be arranged in rows and columns. The first type of the plurality of processing elements_and the second type of the plurality of processing elements_may be arranged in m columns. In addition, the first type of the plurality of processing elements_may be arranged in n rows, and the second type of the plurality of processing elements_may be arranged in 1 rows. Accordingly, the first type of the plurality of processing elements_and the second type of the plurality of processing element_may be arranged in (n+1) rows and m columns.
111 4 500 111 4 1112 111 4 500 The LSU_may receive at least one of data, a control signal, or a synchronization signal from the outside via the local interconnection. The LSU_may transmit at least one of the received data, control signal, or synchronization signal to the CGRA L0 memory. Similarly, the LSU_may transfer at least one of the data, control signal, or synchronization signal to the outside via the local interconnection.
101 101 111 3 111 3 111 3 111 2 111 1 111 4 111 3 111 3 111 2 111 1 111 4 a b a b The neural coremay have a CGRA (Coarse Grained Reconfigurable Architecture) structure. Accordingly, in the neural core, each of the first type of the plurality of processing elements_and the second type of the plurality of processing elements_of the PE array_may be connected to at least one of the CGRA L0 memory_, the instruction memory_, or the LSU_, respectively. In some embodiments, the first type of the plurality of processing elements_and the second type of the plurality of processing elements_do not have to be connected to all of the CGRA L0 memory_, the instruction memory_, and the LSU_, but may be connected to some thereof.
111 3 111 3 111 2 111 1 111 4 111 3 111 3 a b a b Further, the first type of the plurality of processing elements_and the second type of the plurality of processing elements_may be different types of processing elements from each other. Accordingly, out of the CGRA L0 memory_, the instruction memory_, and the LSU_, the elements connected to the first type of the plurality of processing elements_and the elements connected to the second type of the plurality of processing elements_may be different from each other.
101 111 3 111 3 a b The neural coreof the disclosure having a CGRA structure enables high-level parallel calculations, and since direct data exchange between the first type of the plurality of processing elements_and the second type of the plurality of processing elements_is possible, the power consumption may be low. In addition, by including two or more types of processing elements, optimization according to various calculation tasks may also be possible.
111 3 111 3 a b For example, if the first type of the plurality of processing elements_are processing elements that perform two-dimensional calculations, the second type of the plurality of processing elements_may be processing elements that perform one-dimensional calculations. However, the embodiment is not limited thereto.
24 FIG. 25 FIG. is a diagram for illustrating a hierarchical structure of a command processor and task managers of a neural processing device in accordance with some embodiments of the disclosure, andis a diagram for illustrating a hierarchical structure of a command processor and task managers of a neural processing device in accordance with some embodiments of the disclosure.
24 25 FIGS.and 300 7000 300 1 300 300 7000 300 Referring to, if a number of task managersincreases, it may be difficult for the command processorto manage all of the task managers. Therefore, the neural processing devicein accordance with some embodiments of the disclosure may have a hierarchical structure in which each of master task managersM manages the plurality of task managersand the command processormanages the master task managersM.
25 FIG. 300 300 1 300 2 300 1 300 2 300 300 1 300 2 s s s s s s Further, referring to, levels below one of the master task managerM may also be subdivided into a plurality. For example, a first sub-task managerand a second sub-task managermay form each layer. That is, one first sub-task managermay manage at least one second sub-task manager, and one master task managerM may manage at least one first sub-task manager. Additionally, several layers may be added below the second sub-task manageras well.
300 300 7000 300 24 25 FIGS.and That is, although three levels of the task manager, the master task managerM, and the command processorare shown in, the number of levels may be four or more. In some embodiments, the depth of the hierarchical structure may vary as desired depending on the number of task managers.
26 FIG. is a block diagram for illustrating memory reconfiguration of a neural processing system in accordance with some embodiments of the disclosure.
26 FIG. 26 FIG. 10 160 160 a h Referring to, the neural core SoCmay include first to eighth processing unitstoand an on-chip memory OCM. Althoughillustrates eight processing units as an example, this is merely illustrative, and the number of processing units may vary as desired.
120 120 2000 a h The on-chip memory OCM may include first to eighth L0 memoriestoand a shared memory.
120 120 160 160 160 160 120 120 a h a h a h a h The first to eighth L0 memoriestomay be used as private memories for the first to eighth processing unitsto, respectively. In some embodiments, the first to eighth processing unitstoand the first to eighth L0 memoriestomay correspond to each other 1:1.
2000 2100 2100 2100 2100 160 160 120 120 a h a h a h a h The shared memorymay include first to eighth memory unitsto. The first to eighth memory unitstomay correspond to the first to eighth processing unitstoand the first to eighth L0 memoriesto, respectively. That is, the number of memory units may be eight, which is the same as the number of processing units and L0 memories.
2000 2000 2000 The shared memorymay operate in one of two kinds of on-chip memory types. In some embodiments, the shared memorymay operate in one of a L0 memory type or a global memory type. In some embodiments, the shared memorymay implement two types of logical memories with one piece of hardware.
2000 2000 160 160 120 120 2000 a h a h If the shared memoryis implemented in the L0 memory type, the shared memorymay operate as a private memory for each of the first to eighth processing unitsto, just like the first to eighth L0 memoriesto. The L0 memory can operate at a relatively higher clock speed compared with the global memory, and the shared memorymay also use a relatively higher clock speed when operating in the L0 memory type.
2000 2000 160 160 2000 160 160 120 120 a b a h a h. If the shared memoryis implemented in the global memory type, the shared memorymay operate as a common memory used by the first processing unitand the second processing unittogether. In this case, the shared memorymay be shared not only by the first to eighth processing unitstobut also by the first to eighth L0 memoriesto
2000 160 160 2000 2000 32 6000 32 a h 2 FIG. The global memory may generally use a lower clock compared with the L0 memory, but is not limited thereto. When the shared memoryoperates in the global memory type, the first to eighth processing unitstomay share the shared memory. In this case, the shared memorymay be connected to the volatile memoryofvia the global interconnectionand may also operate as a buffer for the volatile memory.
2000 2000 2000 2000 At least part of the shared memorymay operate in the L0 memory type, and the rest may operate in the global memory type. In some embodiments, the entire shared memorymay operate in the L0 memory type, or the entire shared memorymay operate in the global memory type. Alternatively, part of the shared memorymay operate in the L0 memory type, and the rest may operate in the global memory type.
27 FIG. is a block diagram showing an example of memory reconstruction of a neural processing system in accordance with some embodiments of the disclosure.
26 27 FIGS.and 1 3 5 7 160 160 160 160 120 120 120 120 2 4 6 8 160 160 160 160 120 120 120 120 2 4 6 8 2100 2100 2100 2100 2100 2100 2100 2100 2000 a c e g a c e g b d f h b d f h b d f h a c e g With reference to, first, third, fifth, and seventh dedicated areas AE, AE, AE, and AEfor each of the first, third, fifth, and seventh processing units,,, andmay include only the first, third, fifth, and seventh L0 memories,,, and, respectively. Further, second, fourth, sixth, and eighth dedicated areas AE, AE, AE, and AEfor each of the second, fourth, sixth, and eighth processing units,,, andmay include second, fourth, sixth, and eighth L0 memories,,, and, respectively. In addition, the second, fourth, sixth, and eighth dedicated areas AE, AE, AE, and AEmay include the second, fourth, sixth, and eighth memory units,,, and. The first, third, fifth, and seventh memory units,,, andof the shared memorymay be used as a common area AC.
160 160 2 120 2100 2 120 2100 4 6 8 2 a h b b b b The common area AC may be a memory shared by the first to eighth processing unitsto. The second dedicated area AEmay include a second L0 memoryand a second memory unit. The second dedicated area AEmay be an area in which the second L0 memoryand the second memory unitthat are separated hardware-wise operate in the same manner and operate logically as one L0 memory. The fourth, sixth, and eighth dedicated areas AE, AE, and AEmay also operate in the same manner as the second dedicated area AE.
2000 2000 The shared memoryin accordance with the embodiment may convert an area corresponding to each processing unit into a logical L0 memory and a logical global memory of an optimized ratio and may use them. The shared memorymay perform the adjustment of this ratio at runtime.
That is, each processing unit may perform the same task in some cases, but may perform different tasks in other cases as well. In this case, the amount of the L0 memory and the amount of the global memory required for the tasks carried out by each processing unit are inevitably different each time. Accordingly, if the composition ratio of the L0 memory and the shared memory is fixedly set as in the conventional on-chip memory, there may occur inefficiency due to the calculation tasks assigned to each processing unit.
2000 Therefore, the shared memoryof the neural processing device in accordance with the embodiment may set an optimal ratio of the L0 memory and the global memory according to computation tasks during the runtime, and may enhance the efficiency and speed of computation.
28 FIG. 26 FIG. is an enlarged block diagram of a portion A of.
26 28 FIGS.and 2000 122 1 122 1 122 1 122 1 2100 2100 2200 a b e f a h With reference to, the shared memorymay include a first L0 memory controller_, a second L0 memory controller_, a fifth L0 memory controller_, a sixth L0 memory controller_, the first to eighth memory unitsto, and a global controller. Other L0 memory controllers not shown may also be included in the embodiment, but the description thereof will be omitted for convenience.
122 1 122 1 122 1 122 1 2200 122 1 122 1 122 1 122 1 2200 a b e f a b e f The first L0 memory controller_, the second L0 memory controller_, the fifth L0 memory controller_, the sixth L0 memory controller_, and the global controllermay be referred to respectively as a first L0 memory controller circuit, a second L0 memory controller circuit, a fifth L0 memory controller circuit, a sixth L0 memory controller circuit, and a global controller circuit. However, for the sake of convenience, the terms are respectively unified as a first L0 memory controller, a second L0 memory controller, a fifth L0 memory controller, a sixth L0 memory controller, and a global controller. In addition, each of the first L0 memory controller_, the second L0 memory controller_, the fifth L0 memory controller_, the sixth L0 memory controller_, and the global controllermay be implemented as a circuit or circuitry.
122 1 120 122 1 2100 2100 122 1 2100 a a a a a a a. The first L0 memory controller_may control the first L0 memory. In addition, the first L0 memory controller_may control the first memory unit. Specifically, when the first memory unitis implemented in a logical L0 memory type, the control by the first L0 memory controller_may be performed on the first memory unit
122 1 120 122 1 2100 2100 122 1 2100 b b b b b a b. The second L0 memory controller_may control the second L0 memory. Further, the second L0 memory controller_may control the second memory unit. In some embodiments, when the second memory unitis implemented in the logical L0 memory type, the control by the first L0 memory controller_may be performed on the second memory unit
122 1 120 122 1 2100 2100 122 1 2100 e e e e e e e. The fifth L0 memory controller_may control the fifth L0 memory. Further, the fifth L0 memory controller_may control the fifth memory unit. In some embodiments, when the fifth memory unitis implemented in the logical L0 memory type, the control by the fifth L0 memory controller_may be performed on the fifth memory unit
122 1 120 122 1 2100 2100 122 1 2100 f f f f f f f. The sixth L0 memory controller_may control the sixth L0 memory. Further, the sixth L0 memory controller_may control the sixth memory unit. In some embodiments, when the sixth memory unitis implemented in the logical L0 memory type, the control by the sixth L0 memory controller_may be performed on the sixth memory unit
2200 2100 2100 2200 2100 2100 2100 2100 a h a h a h The global controllermay control all of the first to eighth memory unitsto. Specifically, the global controllermay control the first memory unitto the eighth memory unitwhen each of the first to eighth memory unitstooperate logically in the global memory type (i.e., when they do not operate logically in the L0 memory type).
2100 2100 122 1 122 1 2200 a h a h In some embodiments, the first to eighth memory unitstomay be controlled by the first to eighth L0 memory controllers_to_, respectively, or may be controlled by the global controller, depending on what type of memory they are logically implemented.
122 1 122 1 122 1 122 1 2100 2100 122 1 122 1 2100 2100 120 120 160 160 2100 2100 160 160 a b e f a h a h a h a h a h a h a h. If the L0 memory controllers including the first, second, fifth, and sixth L0 memory controllers_,_,_, and_control the first to eighth memory unitsto, respectively, the first to eighth L0 memory controllers_to_control the first to eighth memory unitstoin the same manner as the first to eighth L0 memoriesto, and thus, can control them as the private memory of the first to eighth processing unitsto. Accordingly, the first to eighth memory unitstomay operate at clock frequencies corresponding to the clock frequencies of the first to eighth processing unitsto
122 1 122 1 122 1 122 1 110 a b e f The L0 memory controllers including the first L0 memory controller_, the second L0 memory controller_, the fifth L0 memory controller_, and the sixth L0 memory controller_may each include the LSU.
2200 2100 2100 2200 2100 2100 160 160 2100 2100 160 160 2200 2100 2100 2200 a h a h a h a h a h a h If the global controllercontrols at least one of the first to eighth memory unitsto, respectively, then the global controllermay control the first to eighth memory unitstoas the global memory of the first to eighth processing unitsto, respectively. Accordingly, at least one of the first to eighth memory unitstomay operate at a clock frequency independent of the clock frequencies of the first to eighth processing unitsto, respectively. In some embodiments, if the global controllercontrols the i-th memory unit among the first to eighth memory unitsto, the global controllermay control the i-th memory unit as the global memory of the i-th processing unit, and the i-th memory unit may operate at a clock frequency independent of the clock frequency of the i-th processing unit. However, the embodiment is not limited thereto.
2200 2100 2100 6000 2100 2100 30 2200 120 120 a h a h a h. 3 FIG. 2 FIG. The global controllermay connect the first to eighth memory unitstoto the global interconnectionof. The first to eighth memory unitstomay exchange data with the off-chip memoryofby the control of the global controlleror may respectively exchange data with the first to eighth L0 memoriesto
2100 2100 2100 2110 2110 2100 2110 a h a a a a a 28 FIG. Each of the first to eighth memory unitstomay include at least one memory bank. The first memory unitmay include at least one first memory bank. The first memory banksmay be areas obtained by dividing the first memory unitinto certain sizes. The first memory banksmay all be memory devices of the same size. However, the embodiment is not limited thereto.illustrates that four memory banks are included in one memory unit.
2100 2100 2100 2110 2110 2110 b e f b e f Similarly, the second, fifth, and sixth memory units,, andmay include at least one second, fifth, and sixth memory banks,, and, respectively.
2110 2110 2110 2110 a e b f. In the following, the description will be made based on the first memory banksand the fifth memory banks, which may be the same as other memory banks including the second and sixth memory banksand
2110 2110 2100 a a a Each of the first memory banksmay operate logically in the L0 memory type or operate logically in the global memory type. In this case, the first memory banksmay operate independently of the other memory banks in the first memory unit. However, the embodiment is not limited thereto.
2100 120 120 2100 a a a a. If each memory bank operates independently, the first memory unitmay include a first area operating in the same manner as the first L0 memoryand a second area operating in a different manner from the first L0 memory. In this case, the first area and the second area do not necessarily coexist, but any one area may take up the entire first memory unit
2100 120 120 2100 b b b a. Likewise, the second memory unitmay include a third area operating in the same manner as the second L0 memoryand a fourth area operating in a different manner from the second L0 memory. In this case, the third area and the fourth area do not necessarily coexist, and any one area may take up the entire first memory unit
In this case, the ratio of the first area to the second area may be different from the ratio of the third area to the fourth area. However, the embodiment is not limited thereto. Therefore, the ratio of the first area to the second area may be the same as the ratio of the third area to the fourth area. In some embodiments, the memory composition ratio in each memory unit may vary as desired.
In general, in the case of the conventional system-on-chip, the on-chip memory except for high-speed L0 memory was often composed of high-density, low-power SRAM. This is because SRAM has high efficiency in terms of chip area and power consumption relative to required capacity. However, with the conventional on-chip memory, the processing speed slowed down significantly as was inevitable in the case where tasks that require more data quickly than the predetermined capacity of the L0 memory, and, even when the need for the global memory is not great, there is no way to utilize the remaining global memory, resulting in inefficiency.
2000 2000 On the other hand, the shared memoryin accordance with some embodiments of the disclosure may be controlled selectively by any one of the two controllers depending on the case. In the case depicted, the shared memorymay be controlled not only as a whole by a determined one of the two controllers but also independently for each memory unit or each memory bank.
2000 2000 Through this, the shared memoryin accordance with the embodiment can obtain an optimal memory composition ratio according to calculation tasks during the runtime and can perform faster and more efficient calculation tasks. In the case of a processing unit specialized in artificial intelligence, the required sizes of L0 memory and global memory may vary for each particular application. Moreover, even for the same application, the required sizes of L0 memory and global memory may vary for each layer when a deep learning network is used. In the shared memory, in accordance with the embodiment, the composition ratio of the memory can be changed during runtime even when calculation steps change according to each layer, making fast and efficient deep learning tasks possible.
29 FIG. 28 FIG. 29 FIG. 2110 2110 a a. is a diagram for illustrating the first memory bank ofin detail. Althoughillustrates the first memory bank, other memory banks may also have the same structure as the first memory bank
29 FIG. 2110 1 2 a Referring to, the first memory bankmay include a cell array Ca, a bank controller Bc, a first path unit P, and a second path unit P.
1 2 1 2 In this case, the bank controller Bc, the first path unit P, and the second path unit Pmay be referred to respectively as a bank controller circuit, a first path unit circuit, and a second path unit circuit. However, for the sake of convenience, the terms are respectively unified as a bank controller, a first path unit, and a second path unit. In addition, each of the bank controller Bc, the first path unit P, and the second path unit Pmay be implemented as a circuit or circuitry.
The cell array Ca may include a plurality of memory devices (cells) therein. In the cell array Ca, the plurality of memory devices may be arranged in a lattice structure. The cell array Ca may be, for example, a SRAM (static random-access memory) cell array.
The bank controller Be may control the cell array Ca. The bank controller Be may determine whether the cell array Ca operates in the L0 memory type or in the global memory type, and may control the cell array Ca according to the determined memory type.
1 2 Specifically, the bank controller Be may determine whether to transmit and receive data in the direction of the first path unit Por to transmit and receive data in the direction of the second path unit Pduring the runtime. The bank controller Be may determine a data transmission and reception direction according to a path control signal Spc.
The path control signal Spc may be generated by a pre-designed device driver or compiler. The path control signal Spc may be generated according to the characteristics of calculation tasks. Alternatively, the path control signal Spc may be generated by an input received from a user. In some embodiments, the user may directly apply an input to the path control signal Spc in order to select optimal memory composition ratio.
1 2 The bank controller Be may determine a path along which the data stored in the cell array Ca are transmitted and received via the path control signal Spc. The exchange interface of data may be changed as the bank controller Bc determines the path along which the data are transmitted and received. In some embodiments, a first interface may be used when the bank controller Be exchanges data with the first path unit P, and a second interface may be used when the bank controller Be exchanges data with the second path unit P. In this case, the first interface and the second interface may be different from each other.
Also, address systems in which data are stored may vary as well. In some embodiments, if a particular interface is selected, then read and write operations may be performed in an address system corresponding thereto.
The bank controller Be may operate at a particular clock frequency. For example, if the cell array Ca is an SRAM cell array, the bank controller Be may operate at the operating clock frequency of a general SRAM.
1 1 160 6000 160 120 160 1 2000 1 122 1 122 1 a a a a a b 28 FIG. The first path unit Pmay be connected to the bank controller Bc. The first path unit Pmay directly exchange the data of the cell array Ca with the first processing unit. In this case, “directly” may mean being exchanged with each other without going through the global interconnection. In some embodiments, the first processing unitmay exchange data directly with the first L0 memory, and the first processing unitmay exchange data via the first path unit Pwhen the shared memoryis implemented logically in the L0 memory type. The first path unit Pmay include L0 memory controllers including the first L0 memory controller_and the second L0 memory controller_as shown in.
1 1 160 120 160 160 1 160 a a a a a. The first path unit Pmay form a multi-cycle sync-path. In some embodiments, the operating clock frequency of the first path unit Pmay be the same as the operating clock frequency of the first processing unit. The first L0 memorymay quickly exchange data at the same clock frequency as the operating clock frequency of the first processing unitin order to quickly exchange data at the same speed as the operation of the first processing unit. Likewise, the first path unit Pmay also operate at the same clock frequency as the operating clock frequency of the first processing unit
1 1 In this case, the operating clock frequency of the first path unit Pmay be multiples of the operating clock frequency of the bank controller Bc. In this case, a clock domain crossing (CDC) operation for synchronizing the clocks between the bank controller Be and the first path unit Pis not required separately, and thus, a delay of data transmission may not occur. Accordingly, faster and more efficient data exchange can be possible.
29 FIG. 1 1 1 In the embodiment shown in, an operating clock frequency of the first path unit Pmay be 1.5 GHz, as an example. This may be twice the frequency of 750 MHz of the bank controller Bc. However, the embodiment is not limited thereto, and any operating clock frequency of the first path unit Pmay be possible as long as the first path unit Poperates at integer multiples of the clock frequency of the bank controller Bc.
2 2 160 6000 160 6000 2 160 a a a The second path unit Pmay be connected to the bank controller Bc. The second path unit Pmay exchange the data of the cell array Ca with the first processing unitnot directly but via the global interconnection. In some embodiments, the first processing unitmay exchange data with the cell array Ca via the global interconnectionand the second path unit P. In this case, the cell array Ca may exchange data not only with the first processing unitbut also with other processing units.
2 2110 2 2200 a 28 FIG. In some embodiments, the second path unit Pmay be a data exchange path between the cell array Ca and all the processing units when the first memory bankis implemented logically in the global memory type. The second path unit Pmay include the global controllerof.
2 2 6000 2 6000 The second path unit Pmay form an asynchronous path or Async-Path. The operating clock frequency of the second path unit Pmay be the same as the operating clock frequency of the global interconnection. Likewise, the second path unit Pmay also operate at the same clock frequency as the operating clock frequency of the global interconnection.
22 FIG. 2 2 2 In the case of the embodiment as shown in, the operating clock frequency of the second path unit Pmay not be synchronized with the operating clock frequency of the bank controller Bc. In this case, the clock domain crossing (CDC) operation for synchronizing the clocks between the bank controller Be and the second path unit Pmay be required. If the operating clock frequency of the bank controller Be and the operating clock frequency of the second path unit Pare not synchronized with each other, the degree of freedom in the design of the clock domain may be relatively high. Therefore, the difficulty of hardware design is decreased, thereby making it possible to more easily derive the desired hardware operation.
1 2 1 2 The bank controller Be may use different address systems in the case of exchanging data via the first path unit Pand in the case of exchanging data via the second path unit P. In some embodiments, the bank controller Be may use a first address system if exchanging data via the first path unit Pand a second address system if exchanging data via the second path unit P. In this case, the first address system and the second address system may be different from each other.
A bank controller Be is not necessarily required for each memory bank. In some embodiments, a bank controller Be may not be used to schedule, but instead serves to transfer signals, and thus, is not a required component for each memory bank having two ports. Therefore, one bank controller Be can be operably coupled to control multiple memory banks. The multiple memory banks may operate independently even if they are controlled by the bank controller Bc. However, the embodiment is not limited thereto.
As a matter of course, the bank controller Be may exist for each memory bank. In this case, the bank controller Be may control each memory bank individually.
28 FIG. 29 FIG. 2100 1 2100 2 2100 1 2100 2 a a b b Referring toand, if the first memory unitexchanges data via the first path unit P, the first address system may be used. If the first memory unitexchanges data via the second path unit P, the second address system may be used. Similarly, if the second memory unitexchanges data via the first path unit P, a third address system may be used. If the second memory unitexchanges data via the second path unit P, the second address system may be used. In this case, the first address system and the third address system may be the same as each other. However, the embodiment is not limited thereto.
160 160 160 160 a b a b. The first address system and the third address system may each be used exclusively for the first processing unitand the second processing unit, respectively. The second address system may be commonly applied to the first processing unitand the second processing unit
29 FIG. 2 2 In, the operating clock frequency of the second path unit Pmay operate at 1 GHz, as an example. This may be a frequency that is not synchronized with the operating clock frequency of 750 MHz of the bank controller Bc. In some embodiments, the operating clock frequency of the second path unit Pmay be freely set without being dependent on the operating clock frequency of the bank controller Be at all.
2000 1 2 A generic global memory has used slow SRAM (e.g., 750 MHz) and a global interconnection (e.g., 1 GHz) faster than that, inevitably resulting in delays due to the CDC operation. On the other hand, the shared memoryin accordance with some embodiments has room to use the first path unit Pin addition to the second path unit P, thereby making it possible to avoid delays resulting from the CDC operation.
6000 2000 1 2 2200 Furthermore, in the generic global memory, a plurality of processing units uses one global interconnection, and thus, when an amount of data transfer occurs at the same time, the decrease in the overall processing speed is likely to occur. On the other hand, the shared memoryin accordance with some embodiments has room to use the first path unit Pin addition to the second path unit P, thereby making it possible to achieve the effect of properly distributing the data throughput that could be concentrated on the global controlleras well.
30 FIG. is a block diagram for illustrating a software hierarchy of a neural processing device in accordance with some embodiments.
30 FIG. 10000 20000 30000 Referring to, the software hierarchy of the neural processing device in accordance with some embodiments may include a deep learning (DL) framework, a compiler stack, and a back-end module.
10000 The DL frameworkmay mean a framework for a deep learning model network used by a user. For example, a neural network that has finished training may be generated using a program such as TensorFlow or PyTorch.
20000 21000 22000 23000 24000 25000 The compiler stackmay include an adaptation layer, a compute library, a front-end compiler, a back-end compiler, and a runtime driver.
21000 10000 21000 10000 21000 The adaptation layermay be a layer in contact with the DL framework. The adaptation layermay quantize a neural network model of a user generated by the DL frameworkand modify graphs. In addition, the adaptation layermay convert a type of model into a required type.
23000 21000 24000 The front-end compilermay convert various neural network models and graphs transferred from the adaptation layerinto a constant intermediate representation (IR). The converted IR may be a preset representation that is easy to handle later by the back-end compiler.
23000 23000 The optimization that can be done in advance in the graph level may be performed on such an IR of the front-end compiler. In addition, the front-end compilermay finally generate the IR through the task of converting it into a layout optimized for hardware.
24000 23000 24000 The back-end compileroptimizes the IR converted by the front-end compilerand converts it into a binary file, enabling it to be used by the runtime driver. The back-end compilermay generate an optimized code by dividing a job at a scale that fits the details of hardware.
22000 22000 24000 The compute librarymay store template operations designed in a form suitable for hardware among various operations. The compute libraryprovides the back-end compilerwith multiple template operations required by hardware, allowing the optimized code to be generated.
25000 The runtime drivermay continuously perform monitoring during driving, thereby making it possible to drive the neural network device in accordance with some embodiments. Specifically, it may be responsible for the execution of an interface of the neural network device.
30000 31000 32000 33000 31000 32000 33000 The back-end modulemay include an ASIC (application-specific integrated circuit), an FPGA (field-programmable gate array), and a C-model. The ASICmay refer to a hardware chip determined according to a predetermined design method. The FPGAmay be a programmable hardware chip. The C-modelmay refer to a model implemented by simulating hardware on software.
30000 20000 The back-end modulemay perform various tasks and derive results by using the binary code generated through the compiler stack.
31 FIG. is a conceptual diagram for illustrating deep learning calculations performed by a neural processing device in accordance with some embodiments.
31 FIG. 40000 Referring to, an artificial neural network modelis one example of a machine learning model and is a statistical learning algorithm implemented based on the structure of a biological neural network or is a structure for executing the algorithm, in machine learning technology and cognitive science.
40000 40000 The artificial neural network modelmay represent a machine learning model having an ability to solve problems by learning to reduce the error between an accurate output corresponding to a particular input and an inferred output by repeatedly adjusting the weight of the synapse by nodes. Nodes are artificial neurons that have formed a network by combining synapses, as in a biological neural network. For example, the artificial neural network modelmay include any probabilistic model, neural network model, etc., used in artificial intelligence learning methods such as machine learning and deep learning.
40000 40000 A neural processing device in accordance with some embodiments may implement the form of such an artificial neural network modeland perform calculations. For example, the artificial neural network modelmay receive an input image and may output information on at least a part of an object included in the input image.
40000 40000 40000 41000 40100 44000 40200 42000 43000 41000 44000 41000 44000 44000 42000 43000 31 FIG. The artificial neural network modelmay be implemented by a multilayer perceptron (MLP) including multilayer nodes and connections between them. An artificial neural network modelin accordance with the embodiment may be implemented using one of various artificial neural network model structures including the MLP. As shown in, the artificial neural network modelincludes an input layerthat receives input signals or datafrom the outside, an output layerthat outputs output signals or datacorresponding to the input data, and n (where n is a positive integer) hidden layerstothat are located between the input layerand the output layerand that receive a signal from the input layer, extract characteristics, and forward them to the output layer. Here, the output layerreceives signals from the hidden layerstoand outputs them to the outside.
40000 The learning methods of the artificial neural network modelinclude a supervised learning method for training to be optimized to solve a problem by the input of supervisory signals (correct answers), and an unsupervised learning method that does not require supervisory signals.
40000 41000 44000 40000 41000 42000 43000 44000 40000 40000 The neural processing device may directly generate training data, through simulations, for training the artificial neural network model. In this way, by matching a plurality of input variables and a plurality of output variables corresponding thereto with the input layerand the output layerof the artificial neural network model, respectively, and adjusting the synaptic values between the nodes included in the input layer, the hidden layersto, and the output layer, training may be made to enable a correct output corresponding to a particular input to be extracted. Through such a training phase, it is possible to identify the characteristics hidden in the input variables of the artificial neural network model, and to adjust synaptic values (or weights) between the nodes of the artificial neural network modelso that an error between an output variable calculated based on an input variable and a target output is reduced.
32 FIG. is a conceptual diagram for illustrating training and inference operations of a neural network of a neural processing device in accordance with some embodiments.
32 FIG. Referring to, the training phase may be subjected to a process in which a large number of pieces of training data TD are passed forward to the artificial neural network model NN and are passed backward again. Through this, the weights and biases of each node of the artificial neural network model NN are tuned, and training may be performed so that more and more accurate results can be derived. Through the training phase, the artificial neural network model NN may be converted into a trained neural network model NN_T.
In the inference phase, new data ND may be inputted into the trained neural network model NN_T again. The trained neural network model NN_T may derive result data RD through the weights and biases that have already been used in the training, with the new data ND as input. For such result data RD, what training data TD were used in training and how many pieces of training data TD were used in the training phase may be important.
According to some aspects of the disclosure, a neural processor includes a core array comprising a first neural core and a second neural core, each of which includes controllable port, a memory configured to output data to the core array and receive data as input from the core array and a task manager configured to configure data paths of the core array, wherein the a core array includes a first data line configured to transmit data in a first direction, a second data line configured to transmit data in a second direction opposite to the first direction, a first connection line connected to the a first data line and the first neural core, a second connection line connected to the a second data line and the first neural core, a third connection line connected to the a first data line and the second neural core and a fourth connection line connected to the a second data line and the second neural core, and wherein the first neural core and the second neural core are connected in series via the first data line and the second data line.
According to some aspects, the task manager configures the data paths by controlling on/off of the controllable port included in each of the first neural core and the second neural core.
According to some aspects, the memory includes a first memory connected to the first neural core via the first data line and the second data line, the first neural core comprises a first L0 memory connected to the first connection line and the second connection line, and the second neural core comprises a second L0 memory connected to the third connection line and the fourth connection line.
According to some aspects, the task manager configures a first data path comprising a path through which data are provided from the first memory toward the first L0 memory, from the first L0 memory toward the second L0 memory, and a path through which data are provided from the second L0 memory toward the first L0 memory and from the first L0 memory toward the first memory, by controlling the controllable port.
According to some aspects, the memory further includes a second memory connected to the second neural core via the first data line and the second data line.
According to some aspects, the task manager configures a second data path through which data are provided from the first memory toward the first L0 memory, from the first L0 memory toward the second L0 memory, and from the second L0 memory toward the second memory, and a third data path through which data are provided from the second memory toward the second L0 memory, from the second L0 memory toward the first L0 memory and from the first L0 memory toward the first memory, by controlling the controllable port.
According to some aspects, the task manager configures a fourth data path through which data are provided from the first memory toward the first L0 memory and from the first L0 memory toward the first memory by controlling the controllable port, and configures a fifth data path through which data are provided from the second memory toward the second L0 memory and from the second L0 memory toward the second memory by controlling the controllable port.
According to some aspects, the task manager configures a sixth data path through which data are provided from the first memory toward the first L0 memory and from the first L0 memory toward the second memory by controlling the controllable port.
According to some aspects, the first data line comprises a plurality of data lines, and the task manager turns off some of the plurality of data lines included in the first data line by controlling the controllable port.
According to some aspects, the core array includes a first core array comprising the first neural core and the second neural core and a second core array comprising a third neural core and a fourth neural core that are different from the first neural core and the second neural core, wherein the third neural core and the fourth neural core are connected in series with each other via a third data line and a fourth data line that are different from the first data line and the second data line.
According to some aspects, the neural processor further includes an interconnection through which data are moved, wherein the memory, the first core array, and the second core array are connected to the interconnection, and data are moved between the memory, the first core array, and the second core array via the interconnection.
According to some aspects, the memory includes an L1 shared memory shared by the first neural core and the second neural core.
According to some aspects, the core array is included in a neural core system-on-chip, and the memory includes an off-chip memory external to the neural core system-on-chip.
According to some aspects, the neural processor further includes a first neural processor comprising the core array, a second neural processor that is different from the first neural processor and a shared memory shared by the first neural processor and the second neural processor, wherein the memory comprises the shared memory.
According to some aspects, the controllable ports are implemented by software or firmware.
According to some aspects of the disclosure, a neural processor includes a first neural core, a second neural core connected in series with the first neural core in a first direction, a third neural core connected in series with the second neural core in the first direction, a first memory connected in series with the first neural core in a second direction, which is different from the first direction, and a task manager configured to configure data paths of the first neural core, the second neural core, the third neural core, and the first memory, wherein the data paths include a data movement path in the first direction and a data movement path in the second direction, and the task manager can configure the data paths even when at least one of the first neural core, the second neural core, or the third neural core is inoperative.
According to some aspects, if the first neural core and the third neural core are operative and the second neural core is inoperative, by the task manager, the first neural core is provided with first data from the first memory and generates second data by computing the first data, the second neural core is provided with the second data from the first neural core and provides the second data to the third neural core, and the third neural core is provided with the second data from the second neural core and generates third data by computing the second data.
According to some aspects, data lines connecting the first neural core to the third neural core include a plurality of lines, and the task manager turns on all of the plurality of lines or turns off some of the plurality of lines according to bandwidths of data provided from the first memory.
According to some aspects, the neural processor further includes a fourth neural core that is not directly connected to the first neural core, a fifth neural core connected in series with the fourth neural core in the first direction, and an interconnection configured to perform data exchange between the first neural core and the fourth neural core.
According to some aspects of the disclosure, a neural processor includes a first neural core, a second neural core connected in series with the first neural core in a first direction, a third neural core connected in series with the second neural core in the first direction, a first memory connected in series with the first neural core in a second direction, which is different from the first direction, and a task manager configured to set data paths of the first neural core, the second neural core, the third neural core, and the first memory, wherein the task manager can configure the data paths in real time.
1 10 20 1000 4000 5000 100 7000 300 160 In the present disclosure, the neural processing device, the neural Core SoC, the CPU, the neural processor, the non-volatile memory controller, the volatile memory controller, the neural core, the command processor, the task managerand the processing unitmay be referred to as a processor.
While the inventive concept has been particularly shown and described with reference to exemplary embodiments thereof, it will be understood by those of ordinary skill in the art that various changes in form and details may be made therein without departing from the spirit and scope of the inventive concept as defined by the following claims. It is therefore desired that the embodiments be considered in all respects as illustrative and not restrictive, reference being made to the appended claims rather than the foregoing description to indicate the scope of the disclosure.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
March 19, 2024
July 21, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.