According to one embodiment, a variable length coding device includes circuitry. The circuitry determines N code lengths corresponding to respective N symbols, based on a Huffman tree. In a case where the N code lengths include a first code length longer than a maximum code length, the circuitry selects a first symbol corresponding to the first code length from the N symbols, selects, from the N symbols, a second symbol corresponding to a second code length shorter than the maximum code length, changes the second code length corresponding to the second symbol to a code length obtained by adding one to the second code length, and changes the first code length corresponding to the first symbol to a code length equal to the changed second code length.
Legal claims defining the scope of protection, as filed with the USPTO.
generate a frequency table based on frequencies of occurrence of input symbols for each symbol, the frequency table including N symbols, and N frequencies of occurrence that are associated with the N symbols, respectively; generate a Huffman tree based on the frequency table; determine N code lengths that corresponds to the N symbols, respectively, based on the Huffman tree; select a first symbol corresponding to the first code length from the N symbols; select, from the N symbols, a second symbol corresponding to a second code length that is shorter than the maximum code length; change the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length; and change the first code length corresponding to the first symbol to a code length that is equal to the changed second code length; in a case where the N code lengths include a first code length that is longer than a maximum code length: determine N variable length codes that are assigned to the N symbols, respectively, based on the N code lengths; and convert each of the input symbols into a variable length code, based on the N variable length codes that are assigned to the N symbols, respectively, wherein coding circuitry configured to: N is an integer of two or more, and the variable length code into which each of the input symbols is converted has a bit length between one bit length and the maximum code length inclusive. . A variable length coding device comprising
claim 1 divide L symbols of the N symbols into M symbol sets; generate M representative symbols that represent the M symbol sets, respectively; generate the Huffman tree, based on frequencies of occurrence of respective H symbols and frequencies of occurrence of the respective M representative symbols, the H symbols being obtained by excluding the L symbols from the N symbols; determine H code lengths corresponding to the respective H symbols and M code lengths corresponding to the respective M representative symbols, based on the Huffman tree; change the first representative symbol in order for the first representative symbol to correspond to the fourth code length; and change the third symbol in order for the third symbol to correspond to the third code length; and in a case where a third code length corresponding to a first representative symbol of the M representative symbols is longer than a representative symbol maximum code length and the H symbols include a third symbol corresponding to a fourth code length that is shorter than or equal to the representative symbol maximum code length: select, from the H symbols, the first symbol corresponding to the first code length; select, from the H symbols, the second symbol corresponding to the second code length that is shorter than the maximum code length; change the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length; and change the first code length corresponding to the first symbol to a code length that is equal to the changed second code length, in a case where the H code lengths include the first code length that is longer than the maximum code length: the coding circuitry is configured to: L is an integer of one or more, M is an integer of one or more, and H is an integer of one or more. . The variable length coding device according to, wherein
claim 1 divide L symbols of the N symbols into M symbol sets; generate M representative symbols that represents the M symbol sets, respectively; generate the Huffman tree, based on frequencies of occurrence of respective H symbols and frequencies of occurrence of the respective M representative symbols, the H symbols being obtained by excluding the L symbols from the N symbols; determine H code lengths that correspond to the H symbols, respectively, and M code lengths that correspond to the M representative symbols, respectively, based on the Huffman tree; select, from the H symbols, a third symbol corresponding to a fourth code length that is equal to the third code length; change the third code length corresponding to the first representative symbol to a code length that is obtained by subtracting one from the third code length; and change the fourth code length corresponding to the third symbol to a longer code length; in a case where a third code length corresponding to a first representative symbol of the M representative symbols is longer than a representative symbol maximum code length and the H symbols do not include a symbol corresponding to a code length that is shorter than or equal to the representative symbol maximum code length: select, from the H symbols, two symbols each corresponding to a fifth code length that is equal to a code length obtained by adding one to the changed third code length, further change the changed third code length corresponding to the first representative symbol to a code length that is obtained by subtracting one from the changed third code length; and change the fifth code length corresponding to each of the selected two symbols to a longer code length; in a case where the changed third code length is longer than the representative symbol maximum code length: select, from the H symbols, four symbols each corresponding to a sixth code length that is equal to a code length obtained by adding two to the further changed third code length; change the further changed third code length corresponding to the first representative symbol to a code length obtained by subtracting one from the further changed third code length, and change the sixth code length corresponding to each of the selected four symbols to a longer code length; and in a case where the further changed third code length is longer than the representative symbol maximum code length: select, from the H symbols, the first symbol corresponding to the first code length; select, from the H symbols, the second symbol corresponding to the second code length that is shorter than the maximum code length; change the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length; and change the first code length corresponding to the first symbol to a code length that is equal to the changed second code length, in a case where the H code lengths include the first code length that is longer than the maximum code length: the coding circuitry is configured to: L is an integer of one or more, M is an integer of one or more, and H is an integer of one or more. . The variable length coding device according to, wherein
claim 1 divide L symbols of the N symbols into M symbol sets; generate M representative symbols that represent the M symbol sets, respectively; generate the Huffman tree, based on frequencies of occurrence of respective H symbols and frequencies of occurrence of the respective M representative symbols, the H symbols being obtained by excluding the L symbols from the N symbols; determine H code lengths that correspond to the H symbols, respectively, and M code lengths that correspond to the M representative symbols, respectively, based on the Huffman tree; select an intermediate node whose depth from a root node is shorter than or equal to the representative symbol maximum code length, in the Huffman tree; select a third symbol assigned to each of all leaf nodes that are capable of being reached by tracing from the selected intermediate node in a deeper direction; change a code length corresponding to the third symbol to a longer code length; and change the third code length corresponding to the first representative symbol to the representative symbol maximum code length; and in a case where a third code length corresponding to a first representative symbol of the M representative symbols is longer than a representative symbol maximum code length and the H symbols do not include a symbol corresponding to a code length that is shorter than or equal to the representative symbol maximum code length: select, from the H symbols, the first symbol corresponding to the first code length; select, from the H symbols, the second symbol corresponding to the second code length that is shorter than the maximum code length; change the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length; and change the first code length corresponding to the first symbol to a code length that is equal to the changed second code length, in a case where the H code lengths include the first code length that is longer than the maximum code length: the coding circuitry is configured to: L is an integer of one or more, M is an integer of one or more, and H is an integer of one or more. . The variable length coding device according to, wherein
claim 2 the M symbol sets include a first symbol set that is represented by the first representative symbol, in a case where the first representative symbol is r and a number of symbols that are included in the first symbol set is K(r), the representative symbol maximum code length is obtained by subtracting f(K(r)) from the maximum code length, and f is a function of the number K(r) of the symbols that are included in the first symbol set. . The variable length coding device according to, wherein
claim 5 2 K r ┌log(())┐. the f(K(r)) is defined by . The variable length coding device according to, wherein
claim 1 change code lengths of all symbols each corresponding to a code length that is longer than the maximum code length, among the N symbols, to the maximum code length; calculate a number I of violation symbols among the N symbols; repeat a process as many times as the number I violation symbols, the process including: selecting, from the N symbols, a fourth symbol corresponding to a code length that is equal to the maximum code length; selecting, from the N symbols, a fifth symbol corresponding to a code length that is shorter than the maximum code length; increasing the code length corresponding to the fifth symbol by one; and changing the code length corresponding to the fourth symbol to a code length that is equal to the increased code length. the coding circuitry is configured to: . The variable length coding device according to, wherein
claim 7 the coding circuitry is configured to calculate the number I of violation symbols by . The variable length coding device according to, wherein AllSymbol is a symbol set that includes the N symbols, s is a symbol that is an element of the symbol set AllSymbol, l(s) is a code length corresponding to the symbol s, clip(l(s)) is: the maximum code length if the code length l(s) is longer than the maximum code length; and the code length l(s) if the code length l(s) is shorter than or equal to the maximum code length, and lmax is the maximum code length. where
claim 2 the L symbols are lower L symbols among the N symbols that are sorted in descending order of frequencies of occurrence. . The variable length coding device according to, wherein
claim 2 the L symbols are symbols obtained by excluding upper H symbols among the N symbols that are sorted in descending order of frequencies of occurrence, from the N symbols. . The variable length coding device according to, wherein
claim 2 the M symbol sets include a first symbol set that is represented by the first representative symbol, and the coding circuitry is configured to determine an frequency of occurrence of the first representative symbol, based on frequencies of occurrence of one or more symbols included in the first symbol set. . The variable length coding device according to, wherein
claim 2 the M symbol sets include a first symbol set that is represented by the first representative symbol, the coding circuitry is configured to determine a structure of a subtree that includes leaf nodes to which all symbols in the first symbol set are assigned, respectively, and a code length of a fourth symbol in the first symbol set is obtained by adding a depth from a root node of the subtree to the leaf node to which the fourth symbol is assigned, to a code length corresponding to the first representative symbol. . The variable length coding device according to, wherein
claim 12 the structure of the subtree is a balanced binary tree. . The variable length coding device according to, wherein
claim 12 the coding circuitry is configured to determine the structure of the subtree, based on a frequency of occurrence of each symbol in the first symbol set. . The variable length coding device according to, wherein
claim 12 the coding circuitry is configured to determine the structure of the subtree by selecting one of tree structures prepared in advance. . The variable length coding device according to, wherein
a nonvolatile memory; and claim 1 controller circuitry including the variable length coding device according to, wherein the controller circuitry is configured to write, into the nonvolatile memory, data that includes the variable length code into which each of the input symbols is converted. . A memory system comprising:
generating a frequency table based on frequencies of occurrence of input symbols for each symbol, the frequency table including N symbols, and N frequencies of occurrence that are associated with the N symbols, respectively; generating a Huffman tree based on the frequency table; determining N code lengths that corresponds to the N symbols, respectively, based on the Huffman tree; selecting, from the N symbols, a first symbol corresponding to a first code length, the first code length being included in the N code lengths, the first code length being longer than a maximum code length; selecting, from the N symbols, a second symbol corresponding to a second code length, the second code length being shorter than the maximum code length; changing the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length; changing the first code length corresponding to the first symbol to a code length that is equal to the changed second code length; determining N variable length codes that are assigned to the N symbols, respectively, based on the N code lengths in which the changed first code length and the changed second code length are included; converting each of the input symbols into a variable length code, based on the N variable length codes that are assigned to the N symbols, respectively; and writing, into the nonvolatile memory, data including the variable length code into which each of the input symbols is converted, wherein N is an integer of two or more, and the variable length code into which each of the input symbols is converted has a bit length between one bit length and the maximum code length inclusive. . A method of controlling a nonvolatile memory, the method comprising:
Complete technical specification and implementation details from the patent document.
This application is based upon and claims the benefit of priority from Japanese Patent Application No. 2023-142321, filed Sep. 1, 2023, the entire contents of which are incorporated herein by reference.
Embodiments described herein relate generally to a variable length coding device and a memory system.
Dynamic Huffman coding is a variable length coding for dynamically generating a code table based on a frequency of occurrence of each of symbols to be encoded. The code table indicates correspondence between a symbol and a code word that is assigned to the symbol. In the dynamic Huffman coding, a short code word is assigned to a symbol that occurs at a high frequency, and a long code word is assigned to a symbol that occurs at a low frequency. A code word (variable length code) assigned to a symbol is also referred to as a Huffman code.
In some data compression standards or data compression software (for example, gzip), it may be specified that the length of a code word (code length) assigned to each symbol is restricted not to exceed an upper limit. In such cases, it is necessary to generate a code table such that the length of a code word assigned to each symbol is shorter than or equal to the upper limit at the stage of generating the code table by using a frequency of occurrence of each symbol.
Furthermore, in the data compression standards or data compression software, it may be specified that code words that are assigned to symbols, respectively, in dynamic Huffman coding are perfect codes. The fact that code words that are assigned to symbols, respectively, in dynamic Huffman coding are perfect codes corresponds to the fact that the dynamic Huffman coding is based on a Huffman tree in which every intermediate node has two child nodes. In other words, the fact that code words that are assigned to symbols, respectively, in dynamic Huffman coding are perfect codes means that there is no waste in the code lengths of the assigned code words.
In general, according to one embodiment, a variable length coding device includes coding circuitry. The coding circuitry generates a frequency table based on frequencies of occurrence of input symbols for each symbol. The frequency table includes N symbols, and N frequencies of occurrence that are associated with the N symbols, respectively. The coding circuitry generates a Huffman tree based on the frequency table. The coding circuitry determines N code lengths that corresponds to the N symbols, respectively, based on the Huffman tree. In a case where the N code lengths include a first code length that is longer than a maximum code length, the coding circuitry selects a first symbol corresponding to the first code length from the N symbols, selects, from the N symbols, a second symbol corresponding to a second code length that is shorter than the maximum code length, changes the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length, and changes the first code length corresponding to the first symbol to a code length that is equal to the changed second code length. The coding circuitry determines N variable length codes that are assigned to the N symbols, respectively, based on the N code lengths. The coding circuitry converts each of the input symbols into a variable length code, based on the N variable length codes that are assigned to the N symbols, respectively. N is an integer of two or more. The variable length code into which each of the input symbols is converted has a bit length between one bit length and the maximum code length inclusive.
Various embodiments will be described hereinafter with reference to the accompanying drawings.
1 FIG. 1 2 2 3 shows an example of a configuration of an information processing system that includes a code table generation device according to a first embodiment. The information processing systemincludes a host device(hereinafter referred to as a host) and a memory system.
3 5 3 5 3 3 The memory systemis a semiconductor storage device configured to write data into a nonvolatile memory such as a NAND flash memoryand read data from the nonvolatile memory. The memory systemis implemented as, for example, a solid state drive (SSD) that includes the NAND flash memory. Hereinafter, an example in which the memory systemis implemented as an SSD will be explained, but the memory systemmay be implemented as a hard disk drive (HDD).
2 3 The hostmay be a storage server that stores a large amount of various data in the memory systemor may be a personal computer.
3 2 3 2 2 The memory systemmay be used as a storage for the host. The memory systemmay be provided inside the hostor may be connected to the hostvia a cable or a network.
2 3 An interface for connecting the hostto the memory systemconforms to standards such as SCSI, Serial Attached SCSI (SAS), AT Attachment (ATA), Serial ATA (SATA), PCI Express™ (PCIe™), Ethernet™, Fibre channel, or NVM Express™ (NVMe™).
3 4 5 4 The memory systemincludes a controllerand a NAND flash memory. The controllermay be implemented with circuitry such as a system-on-a-chip (SoC).
3 6 4 6 4 The memory systemmay include a random access memory (RAN) which is a volatile memory, for example, include a dynamic random access memory (DRAM). Alternatively, a RAM such as a static random access memory (SRAM) may be provided inside the controller. Note that the DRAMmay be provided inside the controller.
6 6 The DRAMis a volatile memory. The RAM such as the DRAMincludes, for example, a storage area of firmware (FW) and a cache area of a logical-to-physical address translation table.
5 The NAND flash memoryincludes multiple blocks. Each of the blocks includes multiple pages. The blocks each function as a minimum unit of a data erase operation. A block may be referred to as an erase block or a physical block. Each of the pages includes multiple memory cells connected to a single word line. The pages each function as a unit of a data write operation and a data read operation. Note that a word line may also function as a unit of a data write operation and a data read operation.
The tolerable maximum number of program/erase cycles (maximum number of P/E cycles) for each of the blocks is limited. One P/E cycle of a block includes a data erase operation to erase data stored in all memory cells in the block and a data write operation to write data in each page of the block.
4 11 12 13 14 15 11 12 13 14 15 10 The controllerincludes, for example, a host interface (host I/F), a CPU, a NAND interface (NAND I/F), a DRAM interface (DRAM I/F), and a variable length coding device. The host I/F, the CPU, the NAND I/F, the DRAM I/F, and the variable length coding devicemay be interconnected via a bus.
4 5 13 13 5 The controlleris electrically connected to the NAND flash memoryvia the NAND I/Fthat conforms to an interface standard such as a toggle DDR and an open NAND flash interface (ONFI). The NAND I/Ffunctions as NAND control circuitry configured to control the NAND flash memory.
4 5 The controllerfunctions as a memory controller configured to control the NAND flash memory.
4 5 5 The controllermay function as a flash translation layer (FTL) configured to execute data management and block management of the NAND flash memory. The data management executed by the FTL includes (1) management of mapping information indicative a relationship between each logical address and each physical address of the NAND flash memory, and (2) to hide a difference between data read/data write operations in units of page and data erase operations in units of block. The block management includes management of defective blocks, wear leveling, and garbage collection.
2 3 4 5 5 6 3 The logical address is an address used by the hostfor addressing a storage area of the memory system. Management of mapping between each logical address and each physical address is executed by using a logical-to-physical address translation table. The controlleruses the logical-to-physical address translation table to manage the mapping between each logical address and each physical address with a certain management size. A physical address corresponding to a logical address indicates a physical memory location in the NAND flash memoryto which data of the logical address is written. The logical-to-physical address translation table may be loaded from the NAND flash memoryto the DRAMwhen the memory systemis boot up.
4 4 The data write operation into one page is executable only once in a single P/E cycle. Thus, the controllerwrites updated data corresponding to a logical address not to an original physical memory location in which previous data corresponding to the logical address is stored but to a different physical memory location. Then, the controllerupdates the logical-to-physical address translation table to associate the logical address with this different physical memory location and to invalidate the previous data.
11 3 2 3 11 2 11 2 The host I/Fis a hardware interface that performs communication between the memory systemand the host, which is an external device of the memory system. The host I/Fincludes circuitry which receives various commands, for example, an input/output (I/O) command and a control command, from the host. The I/O commands may be a write command or a read command. The control command may be an unmap command (trim command) or a format command. The host I/Fincludes circuitry that transmits a response to a command and data to the host.
14 6 The DRAM I/Ffunctions as DRAM control circuitry configured to control access to the DRAM.
12 11 13 14 15 12 6 12 12 2 4 The CPUis a processor configured to control the host I/F, the NAND I/F, the DRAM I/F, and the variable length coding device. The CPUperforms various processes by executing the FW loaded to the DRAM. In other words, the FW is a control program for controlling an operation of the CPU. The CPUmay perform, in addition to the above-described processes of FTL, command processes to process various commands received from the host. Note that part of or the entire FTL processes and command processes may be executed by dedicated hardware in the controller.
15 5 12 2 15 15 12 15 The variable length coding deviceis a dynamic Huffman coding unit that encodes data to be written into the NAND flash memoryto compress the data. For example, the CPUinputs write data received from the hostin accordance with receiving a write command, to the variable length coding deviceas plain text data. The variable length coding deviceencodes the plain text data input from the CPU. In order to compress data, the variable length coding devicehas, for example, a configuration for implementing dynamic Huffman coding.
15 The dynamic Huffman coding is a variable length coding for dynamically generating a code table (coding table) by using a frequency of occurrence of each symbol to be encoded. The code table includes information indicative of N types of symbols and N variable length codes (code words) that are associated with the N types of symbols, respectively. N is, for example, an integer of two or more. In the dynamic Huffman coding, a short code word is assigned to a symbol whose frequency of occurrence is high, and a long code word is assigned to a symbol whose frequency of occurrence is low. The variable length coding deviceconverts an input symbol into a code word in accordance with such assignment. Accordingly, a code word obtained by the conversion is a variable length code. Note that the symbol is, for example, data of a fixed length.
A symbol to be encoded is one of the N types of symbols. Hereinafter, a case where N is 256 will be mainly described.
Each of 256 types of symbols is, for example, 1-byte data. In this case, the 256 types of symbols correspond to values from 0 to 255, respectively. Any of the values from 0 to 255 that correspond to the 256 types of symbols, respectively, is also referred to as a symbol number. Note that the number of the types of symbols and the values corresponding to the symbols are mere examples and may be changed in accordance with characteristics of data that includes symbols to be encoded.
15 32 32 32 15 15 The variable length coding deviceincludes a code table generation unit. The code table generation unitis a module that generates a code table for converting a symbol into a variable length code. The code table generation unitmay be a device provided inside the variable length coding device, or may be part of circuitry that realizes the variable length coding device.
32 2 FIG. 11 FIG. Here, generation of a code table in a code table generation unitA according to a comparative example will be described with reference toto.
2 FIG. 32 32 321 322 324 325 326 328 is a block diagram showing a configuration of the code table generation unitA according to the comparative example. The code table generation unitA includes, for example, a frequency table generation unitA, a frequency sorting unitA, a Huffman tree generation unitA, a code length determination unitA, a maximum code length restriction unitA, and a code assignment unitA.
321 40 40 40 321 40 322 The frequency table generation unitA generates a frequency tableA (hereinafter, referred to as a zeroth frequency tableA) by using input symbols. The zeroth frequency tableA is a table indicative of symbols and frequencies of occurrence of the respective symbols. The frequency table generation unitA sends the zeroth frequency tableA to the frequency sorting unitA.
322 40 41 322 41 324 The frequency sorting unitA sorts entries in the zeroth frequency tableA in descending order of the frequencies of occurrence. The frequency table obtained by the sorting is referred to as a first frequency tableA. The frequency sorting unitA sends the first frequency tableA to the Huffman tree generation unitA.
324 41 The Huffman tree generation unitA generates a Huffman tree by using the first frequency tableA.
3 FIG. 324 is a flowchart showing an example of the procedure of a Huffman tree generation process executed by the Huffman tree generation unitA.
324 41 11 324 41 324 First, the Huffman tree generation unitA adds all symbols in the first frequency tableA that have frequencies of occurrence higher than 0, as leaf nodes on a Huffman tree (step S). In other words, the Huffman tree generation unitA generates a Huffman tree including as many leaf nodes as the symbols in the first frequency tableA that have frequencies of occurrence higher than 0. The number of symbols for which the Huffman tree is to be generated (i.e., the number of leaf nodes of the Huffman tree to be generated) by the Huffman tree generation unitA is at most N.
324 12 324 13 324 14 Next, the Huffman tree generation unitA selects a node A having the lowest frequency of occurrence and a node B having the second lowest frequency of occurrence from all leaf and intermediate nodes having no parent nodes in the Huffman tree (step S). The Huffman tree generation unitA adds an intermediate node having the selected nodes A and B as children to the Huffman tree (step S). Then, the Huffman tree generation unitA sets the sum of the frequencies of occurrence of the nodes A and B as a frequency of occurrence of the added intermediate node (step S).
324 15 15 324 12 324 The Huffman tree generation unitA determines whether or not there are two or more leaf and intermediate nodes having no parent nodes in all in the Huffman tree (step S). When there are two or more leaf and intermediate nodes having no parent nodes in all in the Huffman tree (Yes in step S), the process by the Huffman tree generation unitA proceeds to step S. That is, the Huffman tree generation unitA further performs a procedure for adding an intermediate node that has leaf and intermediate nodes having no parent nodes as children.
15 324 When the number of leaf and intermediate nodes having no parent nodes in the Huffman tree is smaller than two (No in step S), the Huffman tree generation unitA ends the Huffman tree generation process.
324 324 With the above Huffman tree generation process, the Huffman tree generation unitA can generate the Huffman tree. The generated Huffman tree includes the leaf nodes respectively corresponding to the symbols whose frequencies of occurrence are higher than 0. The Huffman tree generation unitA constructs the Huffman tree bottom-up from a leaf node corresponding to a symbol that has a lower frequency of occurrence.
2 FIG. 324 325 The description returns to. The Huffman tree generation unitA sends the generated Huffman tree to the code length determination unitA.
325 324 325 325 326 The code length determination unitA determines a code length of each of the symbols by using the Huffman tree received from the Huffman tree generation unitA. A depth of a leaf node starting from a root node (i.e., the number of edges traced from the root node to the leaf node) corresponds to a code length of the corresponding symbol. As the distance from the root node to a node is longer, the depth of the node is deeper and a code length of a symbol corresponding to the node is longer. Thus, the code length determination unitA can determine the code length of each of the symbols by using the Huffman tree. The code length determination unitA sends the determined code length of each of the symbols to the maximum code length restriction unitA.
326 326 325 326 328 7 FIG. 11 FIG. In some data compression standards, code lengths may be restricted not to exceed an upper limit (that is, limited to the upper limit). In such a case, the maximum code length restriction unitA executes a process of restricting (limiting) the code length to the upper limit (hereinafter referred to as a maximum code length restriction process). Specifically, the maximum code length restriction unitA restricts the code length of each of the symbols determined by the code length determination unitA to a code length that is shorter than or equal to the upper limit. The maximum code length restriction unitA sends the code length of each of the symbols restricted to the code length that is shorter than or equal to the upper limit, to the code assignment unitA. The upper limit of code length is also referred to as a maximum code length in the following descriptions. A specific example of the maximum code length restriction process will be described later with reference toto.
328 326 328 26 FIG. The code assignment unitA generates a code table by using the code length of each of the symbols received from the maximum code length restriction unitA. The code assignment unitA generates the code table according to, for example, canonical Huffman coding. The canonical Huffman coding is capable of determining a code bit string (variable length code) to be assigned to a symbol, using only the code lengths of the symbols. A specific example of generating a code table according to the canonical Huffman coding will be described later with reference to.
41 4 FIG. 6 FIG. A specific example in which the code lengths are determined by using the first frequency tableA will be described with reference toto.
4 FIG. 41 321 322 41 shows an example of a configuration of the first frequency tableA generated by the frequency table generation unitA and then sorted by the frequency sorting unit. The first frequency tableA includes multiple entries that correspond to symbols, respectively. Each of the entries includes, for example, a symbol field and a frequency of occurrence field.
The symbol field indicates a corresponding symbol.
The frequency of occurrence field indicates a frequency of occurrence of the corresponding symbol. More specifically, the frequency of occurrence field indicates, for example, the number of times the corresponding symbol occurs in the input data to be processed.
41 Note that, hereinafter, a value indicated in the symbol field is also simply referred to as a symbol. The same will apply to values indicated in other fields of the first frequency tableA and values indicated in fields of other tables.
4 FIG. In the example shown in, a frequency of occurrence of a symbol “a” is 120. A frequency of occurrence of a symbol “b” is 60. A frequency of occurrence of a symbol “c” is 29. A frequency of occurrence of a symbol “d” is 14. A frequency of occurrence of a symbol “e” is four. A frequency of occurrence of each of symbols “f” and “g” is three. A frequency of occurrence of a symbol “h” is two. A frequency of occurrence of each of symbols “i” and “j” is 0.
5 FIG. 4 FIG. 50 324 324 50 41 shows an example of a Huffman treegenerated by the Huffman tree generation unitA. Here, a case where the Huffman tree generation unitA generates the Huffman treeby using the first frequency tableA illustrated inwill be explained.
324 41 324 50 511 521 531 541 561 562 563 564 First, the Huffman tree generation unitA selects, from the first frequency tableA, the eight symbols “a”, “b”, “c”, “d”, “e”, “f”, “g”, and “h” that have the frequencies of occurrence larger than 0. The Huffman tree generation unitA generates the Huffman treeincluding eight leaf nodes,,,,,,, andto which the selected eight symbols are assigned, respectively. For each of the leaf nodes, the frequency of occurrence of the corresponding symbol is set.
324 564 563 50 324 552 564 563 50 324 564 563 552 The Huffman tree generation unitA selects the leaf nodeof the symbol “h” having the smallest frequency of occurrence and the leaf nodeof the symbol “g” having the next smallest frequency of occurrence from all leaf and intermediate nodes in the Huffman treethat have no parent nodes. Then, the Huffman tree generation unitA adds an intermediate nodehaving the selected leaf nodesandas children to the Huffman tree. The Huffman tree generation unitA sets the sum of the frequency of occurrence of the leaf nodeand the frequency of occurrence of the leaf node(=2+3=5) as the frequency of occurrence of the intermediate node.
324 562 561 50 324 551 562 561 50 324 562 561 551 Next, the Huffman tree generation unitA selects the leaf nodeof the symbol “f” having the smallest frequency of occurrence and the leaf nodeof the symbol “e” having the next smallest frequency of occurrence, from all leaf and intermediate nodes in the Huffman treethat have no parent nodes. Then, the Huffman tree generation unitA adds an intermediate nodehaving the selected leaf nodesandas children to the Huffman tree. The Huffman tree generation unitA sets the sum of the frequency of occurrence of the leaf nodeand the frequency of occurrence of the leaf node(=3+4=7) as the frequency of occurrence of the added intermediate node.
324 552 551 50 324 542 552 551 50 324 552 551 542 Next, the Huffman tree generation unitA selects the intermediate nodehaving the smallest frequency of occurrence and the intermediate nodehaving the next smallest frequency of occurrence, from all leaf and intermediate nodes in the Huffman treethat have no parent nodes. Then, the Huffman tree generation unitA adds an intermediate nodehaving the selected intermediate nodesandas children to the Huffman tree. The Huffman tree generation unitA sets the sum of the frequency of occurrence of the intermediate nodeand the frequency of occurrence of the intermediate node(=5+7=12) as the frequency of occurrence of the added intermediate node.
324 50 501 50 5 FIG. Similarly, the Huffman tree generation unitA repeats the operation of adding an intermediate node until the number of leaf and intermediate nodes in the Huffman treethat have no parent nodes becomes one or less (i.e., until only a root nodeis obtained as a node having no parent node). As a result, the Huffman treeshown incan be constructed.
325 50 325 50 501 511 501 564 501 The code length determination unitA determines code lengths that are associated with the symbols, respectively, by using the constructed Huffman tree. Specifically, the code length determination unitA determines the code length of each of the eight symbols “a”, “b”, “c”, “d”, “e”, “f”, “g”, and “h” by using the constructed Huffman tree. The depth of a leaf node corresponding to a symbol starting from the root nodeindicates the code length of the symbol. For example, the depth of the leaf nodeof the symbol “a” starting from the root nodeis one. Thus, the code length of the symbol “a” is 1 bit. In addition, for example, the depth of the leaf nodeof the symbol “h” starting from the root nodeis six. Thus, the code length of the symbol “h” is 6 bits.
6 FIG. 5 FIG. 325 50 shows an example of the code length of each symbol determined by the code length determination unitA by using the Huffman treeshown in.
Specifically, the code length of the symbol “a” is 1 bit. The code length of the symbol “b” is 2 bits. The code length of the symbol “c” is 3 bits. The code length of the symbol “d” is 4 bits. In addition, the code length of each of the symbols “e”, “f”, “g”, and “h” is 6 bits.
7 FIG. 326 325 is a flowchart showing an example of the procedure of the maximum code length restriction process executed by the maximum code length restriction unitA. The maximum code length restriction process is a process of restricting the code lengths of the respective symbols determined by the code length determination unitA to the maximum code length (i.e., to a code length shorter than or equal to the maximum code length).
326 21 326 22 326 23 First, the maximum code length restriction unitA changes all the code lengths of the symbols that are longer than the maximum code length, to the maximum code length (step S). The maximum code length restriction unitA selects P symbols whose code lengths are shorter than the maximum code length, in the order of longer code lengths (step S). P is an integer of one or more. Then, the maximum code length restriction unitA changes the code lengths of the selected P symbols to the maximum code lengths in parallel (step S). As P is larger, the number of symbols whose code lengths are changed to the maximum code lengths at a time is increased. In the following descriptions, P is also referred to as a degree of parallelism.
326 24 Next, the maximum code length restriction unitA calculates a value K of the left-hand side of Kraft's inequality by using the code lengths of all the symbols including the changed code lengths (step S). The Kraft's inequality is expressed by using the following equation (1).
Note that AllSymbol is a symbol set including all the symbols. s is an element (i.e., a symbol) of the symbol set AllSymbol. l(s) is a code length corresponding to symbol s.
When code length l(s) of symbol s∈AllSymbol satisfies Kraft's inequality, existence of a valid code table is guaranteed. In other words, encoding and decoding that use the code table generated based on code length l(s) of symbol s∈AllSymbol can be correctly executed.
In contrast, when code length l(s) of symbol s∈AllSymbol does not satisfy Kraft's inequality, there is no valid code table. Therefore, encoding and decoding that use a code table generated based on code length l(s) of symbol s∈AllSymbol cannot be correctly executed.
326 25 326 The maximum code length restriction unitA determines whether or not the calculated value K of the left-hand side of Kraft's inequality is smaller than or equal to one (step S). In other words, the maximum code length restriction unitA determines whether or not code length l(s) of symbol s E AllSymbol satisfies Kraft's inequality.
25 326 22 326 When the value K of the left-hand side of Kraft's inequality is larger than one (No in step S), the process by the maximum code length restriction unitA returns to step S. In other words, the maximum code length restriction unitA repeats a process of: selecting P symbols whose code lengths are shorter than the maximum code length and that have not yet been selected, in the order of longer code lengths; changing the code lengths of the selected P symbols to the maximum code length; and calculating the value K of the left-hand side of Kraft's inequality, until code length l(s) of symbol s E AllSymbol satisfies Kraft's inequality (i.e., until the value of the left-hand side of Kraft's inequality is smaller than or equal to one).
25 326 When the value K of the left-hand side of Kraft's inequality is smaller than or equal to one (Yes in step S), the maximum code length restriction unitA ends the maximum code length restriction process.
326 326 326 With the above maximum code length restriction process, the maximum code length restriction unitA can change a code length associated with a symbol that is longer than the maximum code length, to the maximum code length. The code lengths of all the symbols including the changed code length do not satisfy Kraft's inequality. Therefore, the maximum code length restriction unitA further performs a process of selecting P symbols each corresponding to a code length shorter than the maximum code length and changing the code lengths of the selected P symbols to the maximum code length, repeatedly until Kraft's inequality is satisfied. The maximum code length restriction unitA can thereby associate all the symbols with code lengths that are shorter than or equal to the maximum code length and satisfy the Kraft's inequality.
8 FIG. 10 FIG. 5 FIG. 5 FIG. 5 FIG. 326 50 50 324 326 50 With reference toto, an operation of the maximum code length restriction unitA to restrict the code lengths of the symbols to the maximum code length will be explained by using modified examples of the Huffman treeof. The Huffman treeofis a Huffman tree constructed by the Huffman tree generation unitA, based on the frequencies of occurrence of the symbols, without considering the maximum code length. The maximum code length restriction unitA transforms the Huffman treeofto satisfy both the restriction on the maximum code length and Kraft's inequality. Here, it is assumed that the maximum code length is 4 bits. In addition, the degree of parallelism P is one in order to make the description easy to understand.
8 FIG. 326 shows an example in which the maximum code length restriction unitA determines symbols each having a code length longer than the maximum code length and a symbol whose code length is to be lengthened to the maximum code length.
326 561 562 563 564 326 Specifically, the maximum code length restriction unitA determines the four symbols “e”, “f”, “g”, and “h” (i.e., leaf nodes,,, and) each having the code length longer than 4 bits. The maximum code length restriction unitA changes the code lengths of the determined four symbols “e”, “f”, “g”, and “h” from 6 bits to 4 bits.
326 531 326 In addition, the maximum code length restriction unitA determines the symbol “c” (i.e., leaf node) having the longest code length among the symbols each having the code length shorter than 4 bits. The maximum code length restriction unitA changes the code length of the determined symbol “c” from 3 bits to 4 bits.
9 FIG. 50 shows the Huffman treein which the code lengths of the four symbols “e”, “f”, “g”, and “h” and the code length of the symbol “c” have been changed according to the above-described changes.
531 543 544 531 50 543 544 In accordance with the code length of the symbol “c” being changed from 3 bits to 4 bits, the nodeis changed from a leaf node to an intermediate node, and two leaf nodesandhaving the intermediate nodeas their parent node are added to the Huffman tree. The symbol “c” is assigned to the leaf node. The symbol “e” is assigned to the leaf node.
50 541 545 546 542 532 50 50 In addition, in the Huffman tree, in accordance with the code lengths of the four symbols “e”, “f”, “g”, and “h being changed from 6 bits to 4 bits, four leaf nodes,,, andhaving an intermediate nodeas their parent node are generated. Accordingly, this Huffman treeis an invalid Huffman tree. Therefore, Huffman codes determined based on this Huffman treeare invalid Huffman codes that do not satisfy Kraft's inequality (i.e., K>1).
326 326 521 326 Thus, the maximum code length restriction unitA further determines a leaf node whose code length is to be lengthened to the maximum code length. Specifically, the maximum code length restriction unitA determines the symbol “b” (i.e., leaf node) having the longest code length among the symbols whose code lengths are shorter than 4 bits. The maximum code length restriction unitA changes the code length of the determined symbol “b” from 2 bits to 4 bits.
10 FIG. 50 shows the Huffman treein which the code length of the symbol “b” has been changed.
521 533 534 521 547 548 533 545 546 534 50 50 50 50 In accordance with the code length of the symbol “b” being changed from 2 bits to 4 bits, the nodeis changed from a leaf node to an intermediate node, and two intermediate nodesandhaving the intermediate nodeas their parent node, two leaf nodesandhaving the intermediate nodeas their parent node, and two leaf nodesandhaving the intermediate nodeas their parent node are added to the Huffman tree. This Huffman treeis a valid Huffman tree. Therefore, Huffman codes determined based on this Huffman treeare valid Huffman codes that satisfy Kraft's inequality (i.e., K<1).
50 326 548 326 548 The Huffman treeobtained by the maximum code length restriction unitA may include the redundant leaf nodeto which no symbol is assigned. In this case, the maximum code length restriction unitA further removes the redundant leaf nodeto which no symbol is assigned.
11 FIG. 50 548 shows the Huffman treein which the redundant leaf nodehas been removed.
548 547 533 533 547 533 A binary tree that includes the redundant leaf node, the leaf node, and the intermediate nodeis merged to become the leaf node. The process of merging a binary tree having a redundant leaf node is also simply referred to as a merge process. The symbol “b” assigned to the leaf nodeis assigned to the leaf node. In other words, the code length of the symbol “b” is reduced by one bit, thereby changed from 4 bits to 3 bits.
50 50 50 50 This Huffman treeis a valid Huffman tree. Therefore, Huffman codes determined based on this Huffman treeare valid Huffman codes that satisfy Kraft's inequality and that have no redundant code assignment (i.e., K=1). In other words, the Huffman codes determined based on the Huffman treeare perfect codes since every intermediate node of the Huffman treehas two child nodes.
32 50 326 32 7 FIG. Thus, in the code table generation unitA of the comparative example, the process of constructing the Huffman treeto determine the code length of each symbol and restricting the code lengths to the maximum code length may require the merge process. There is a case where this merge process may be necessary not only to improve a coding efficiency of Huffman codes, but also to satisfy a requirement that the Huffman codes are perfect codes. This case is, for example, a case where compression standards or compression software require that Huffman codes are perfect codes. When the maximum code length restriction process shown inhas been completed, the maximum code length restriction unitA cannot guarantee that the Huffman codes assigned based on the code length of each symbol are perfect codes, and may have to additionally execute the merge process. For this reason, the code table generation unitA of the comparative example may require additional processing time and additional processing resources (for example, hardware cost).
15 32 15 In contrast, the variable length coding encoderaccording to the embodiment realizes a maximum code length restriction process that guarantees that Huffman codes are perfect codes. In the maximum code length restriction process, the code table generation unitof the variable length coding devicerestricts code lengths of symbols to the maximum code length (that is, to a code length shorter than or equal to the maximum code length) and guarantees that Huffman codes assigned based on the code length of each symbol are perfect codes.
32 32 32 32 32 For example, in a case where N code lengths corresponding to respective N symbols determined based on a Huffman tree include a first code length longer than the maximum code length, the code table generation unitrestricts the first code length to the maximum code length in the maximum code length restriction process. Specifically, the code table generation unitselects at least one first symbol corresponding to the first code length from the N symbols. The code table generation unitselects at least one second symbol corresponding to a second code length that is shorter than the maximum code length, from the N symbols. The code table generation unitchanges the second code length corresponding to the second symbol to a code length obtained by adding one to the second code length. Then, the code table generation unitchanges the first code length corresponding to the first symbol to a code length equal to the changed second code length.
32 32 In other words, the code table generation unitselects the first symbol that does not satisfy a code length restriction of being the code length shorter than or equal to the maximum code length, and selects another second symbol that satisfies the code length restriction. Then, the code table generation unitsets a leaf node on the Huffman tree corresponding to the second symbol as an intermediate node, and then connects leaf nodes to which the first and second symbols are assigned, respectively, as children nodes of the intermediate node. Accordingly, the process for satisfying the code length restriction can be executed while the Huffman codes assigned based on the Huffman tree being perfect codes are maintained.
32 32 32 Therefore, the code table generation unitcan guarantee that the Huffman codes assigned based on the code length of each symbol are perfect codes when the maximum code length restriction process has been completed. Therefore, the code table generation unitof the embodiment can reduce the processing time and the processing resources as compared to the code table generation unitA of the comparative example, which additionally executes the merge process.
12 FIG. 15 15 31 32 33 shows an example of a configuration of the variable length coding deviceaccording to the embodiment. The variable length coding deviceincludes, for example, a buffer unit, the code table generation unit, and a variable length coding unit.
31 15 31 33 32 The buffer unitstores (buffers) symbols input to the variable length coding device. The buffer unitdelays the stored symbols until, for example, a specific timing and then sends the symbols to the variable length coding unit. The specific timing is, for example, the timing when the code table generation unithas completed generation of a code table.
32 15 The code table generation unitgenerates a code table by using the symbols input to the variable length coding device. The code table includes information indicative of symbols and variable length codes (i.e., code bit strings) that are associated with the symbols, respectively.
32 32 32 33 More specifically, the code table generation unitgenerates the code table, based on frequencies of occurrence of symbols that are included in input data of a specific unit. The specific unit may be a unit of a specific data amount or may be a specific group such as a file. In a case where the unit is a specific group, the code table generation unitdetects data indicative of the end of the input data to recognize the input data of the specific unit. The code table generation unitsends the generated code table to the variable length coding unit.
33 31 32 33 The variable length coding unitconverts the symbols sent from the buffer unitinto variable length codes (code bit strings) by using the code table sent from the code table generation unit. The variable length coding unitoutputs the variable length codes obtained by the conversion as compressed data.
15 5 2 12 5 13 With the above configuration, the variable length coding unitcan perform dynamic Huffman coding on the input symbols to convert the input symbols into variable length codes. For example, in a case where the input symbols are data required to be written into the NAND flash memoryby the host, the CPUwrites compressed data that includes one or more variable length codes and the code table that is compressed, into the NAND flash memoryvia the NAND I/F.
4 33 12 5 13 12 15 5 13 2 11 12 5 13 12 12 2 2 2 12 5 2 The controllermay further include an ECC encoder and an ECC decoder. In this case, the ECC encoder generates a parity for error correction (ECC parity) for the compressed data output from the variable length coding unitand generates a code word having the generated ECC parity and the compressed data. The CPUis configured to write the code word to the NAND flash memoryvia the NAND I/F. In other words, the CPUis configured to write data based on the compressed data output from the variable length coding deviceinto the NAND flash memoryvia the NAND I/F. Further, for example, in a case where a read command is received from the hostvia the host I/F, the CPUreads the data based on the read command from the NAND flash memoryvia the NAND I/F. The ECC decoder executes an error correction process on the read data. The read data on which the error correction process has been executed is input to a decompressor by the CPUas compressed data, and the decompressor decompresses the input compressed data. The CPUtransmits the decompressed data to the hostin response to the read command from the host. In other words, in response to the read command from the host, the CPUis configured to decompress data based on data read from the NAND flash memoryand transmit the decompressed data to the host.
15 Note that part or all of the variable length coding devicemay be implemented as hardware such as circuitry or implemented as programs (i.e., software) executed by at least one processor.
32 32 321 322 323 324 325 326 327 328 Next, a specific configuration of the code table generation unitwill be described. The code table generation unitincludes, for example, a frequency table generation unit, a frequency sorting unit, a symbol merge unit, a Huffman tree generation unit, a code length determination unit, a maximum code length restriction unit, a representative symbol expansion unit, and a code assignment unit.
321 40 40 321 40 321 40 40 321 40 322 The frequency table generation unitgenerates a frequency table(hereinafter referred to as a zeroth frequency table) based on frequencies of occurrence of the input symbols for each symbol. For example, the frequency table generation unitcounts the number of occurrences of the input symbols for each symbol, thereby generating the zeroth frequency table. For example, the frequency table generation unitgenerates the zeroth frequency tableevery time 4096 symbols are input. The zeroth frequency tableis a table indicative of symbols and frequencies of occurrence (e.g., the numbers of times of occurrence) that are associated with the symbols, respectively. The frequency table generation unitsends the zeroth frequency tableto the frequency sorting unit.
322 40 41 322 41 323 The frequency sorting unitsorts the entries in the zeroth frequency tablein descending order of the frequencies of occurrence. The frequency table obtained by the sorting is referred to as a first frequency table. The frequency sorting unitsends the first frequency tableto the symbol merge unit.
13 FIG. 13 FIG. 40 321 40 shows an example of the zeroth frequency tablegenerated by the frequency table generation unit. The zeroth frequency tableincludes N entries that correspond to N types of symbols, respectively. Indexes from 0 to N−1 are assigned to the N entries, respectively, in order from the head. Thus, each of the N entries is identifiable by the index. In the example shown in, N is 256. Each entry includes a symbol number field and a frequency of occurrence field.
In an entry corresponding to a symbol, the symbol number field indicates a symbol number corresponding to the symbol. The frequency of occurrence field indicates a frequency (for example, the number of times) at which the corresponding symbol occurs in one or more symbols included in the input data.
Thus, each entry indicates a combination of a symbol number and a frequency of occurrence. The combination of a symbol number and a frequency of occurrence is also referred to as a <symbol number, frequency of occurrence> pair.
40 13 FIG. In the zeroth frequency tableshown in, for example, an entry with the index “0” indicates that a frequency of occurrence of a symbol with a symbol number “0” is three. For example, an entry with the index “1” indicates that a frequency of occurrence of a symbol with a symbol number “1” is 16. Furthermore, for example, an entry with the index “253” indicates that a frequency of occurrence of a symbol with a symbol number “253” is 30.
14 FIG. 41 322 41 40 41 shows an example of the first frequency tableacquired by the frequency sorting unit. The first frequency tableis a table in which the 256 entries in the zeroth frequency tableare arranged in descending order of the frequencies of occurrence. Indexes from 0 to 255 are assigned to the 256 entries in the first frequency table, respectively, in order from the head.
41 14 FIG. In the first frequency tableshown in, for example, an entry with the index “0” indicates that a frequency of occurrence of a symbol with a symbol number “64” is 50. For example, an entry with the index 1 indicates that the frequency of occurrence of the symbol with the symbol number “253” is 30. Furthermore, for example, an entry with the index 255 indicates that a frequency of occurrence of a symbol with a symbol number “8” is 0.
41 As described above, in the first frequency table, the 256 entries are arranged in descending order of the frequencies of occurrence.
12 FIG. The description returns to.
323 41 323 323 41 323 351 352 353 354 The symbol merge unitdivides L symbols among all the symbols (i.e., N symbols) included in the first frequency tableinto M symbol sets, and performs a process for regarding each of the M symbol sets as one symbol (hereinafter referred to as a representative symbol) while a Huffman tree is constructed. The symbol merge unitgenerates M representative symbols that represents the M symbol sets, respectively. L is an integer of one or more. M is an integer of one or more. Hereinafter, a case where M is one will be mainly described in order to make the description easy to understand. In this case, the symbol merge unitperforms a process of regarding the L symbols among all the symbols included in the first frequency tableas one representative symbol. Note that in a case where M is two or more, a process for the one representative symbol to be described below is executed for each of M representative symbols. The symbol merge unitincludes, for example, a symbol distribution unit, a representative symbol frequency estimation unit, a merge symbol count unit, and a merge symbol additional code length determination unit.
351 41 431 432 431 41 432 41 The symbol distribution unitdivides the first frequency tableinto a first part tableand a second part table. The first part tableincludes, for example, the top H entries in descending order of the frequencies of occurrence among the N entries included in the first frequency table. H is an integer of one or more. In addition, the sum of H and L is N. The H entries are entries of H symbols having high frequencies of occurrence. The H symbols having high frequencies of occurrence are also referred to as high-order symbols. The second part tableincludes, for example, the remaining L (=N−H) entries that are obtained by excluding the top H entries from the N entries included in the first frequency table. The L entries are entries of L symbols having low frequencies of occurrence. The L entries are more likely to include a symbol whose frequency of occurrence is 0 than the high-order H entries. The L symbols having the low frequencies of occurrence are also referred to as low-order symbols. The L low-order symbols are regarded as one representative symbol while a Huffman tree is constructed. In other words, the representative symbol is a symbol that represents the L low-order symbols. In contrast, each of the H high-order symbols represents its own high-order symbol and does not represent any other symbol. For this reason, the high-order symbols are also referred to as non-representative symbols.
351 432 41 431 351 431 41 432 432 431 Note that, for example, the symbol distribution unitdecides to include, in the second part table, the bottom (lower-order) L entries in descending order of the frequencies of occurrence among the N entries included in the first frequency table, and then decides to include, in the first part table, the remaining entries that are obtained by excluding the L entries from the N entries. Alternatively, the symbol distribution unitmay decide to include, in the first part table, the top (higher-order or upper-order) H entries in descending order of the frequencies of occurrence among the N entries included in the first frequency table, and then decide to include, in the second part table, the remaining entries obtained by excluding the H entries from the N entries. In other words, which of the entries included in the second part tableand the entries included in the first part tableis preferentially decided can be freely determined.
351 431 324 351 432 352 353 354 The symbol distribution unitsends the first part tableto the Huffman tree generation unit. The symbol distribution unitsends the second part tableto the representative symbol frequency estimation unit, the merge symbol count unit, and the merge symbol additional code length determination unit.
15 FIG. 15 FIG. 41 431 432 41 is a diagram showing (A) an example of the first frequency table, and (B) an example of the first part tableand (C) the second part tablethat are obtained by dividing the first frequency table. In the example shown in, N is 256, H is 32, and L is 224.
41 41 41 431 432 15 FIG.(A) 14 FIG. The first frequency tableshown inis a table in which the N entries are arranged in descending order of the frequencies of occurrence, similarly to the first frequency tableshown in. The first frequency tableis divided into the first part tableand the second part table.
15 FIG.(B) 431 41 431 As shown in, the first part tableincludes the high-order H entries (i.e., H entries from the head) among the entries included in the first frequency table. Indexes from 0 to H−1 are respectively assigned to the H entries in the first part tablein order from the head.
15 FIG.(C) 432 41 432 As shown in, the second part tableincludes the low-order L entries (i.e., L entries from the bottom) among the entries included in the first frequency table. Indexes from 0 to L−1 are respectively assigned to the L entries in the second part tablein order from the head.
12 FIG. The description returns to.
432 352 352 352 In order that the L low-order symbols included in the second part tableare regarded as one representative symbol while a Huffman tree is constructed, the representative symbol frequency estimation unitdetermines a frequency of occurrence of the representative symbol. The representative symbol frequency estimation unitdetermines the frequency of occurrence of the representative symbol, based on the frequencies of occurrence of the L low-order symbols. Specifically, the representative symbol frequency estimation unitestimates the frequency of occurrence of the representative symbol by, for example, estimating the sum of the frequencies of occurrence of the L low-order symbols.
352 432 Here, an example of calculating an estimated value of the sum of frequencies of occurrence of 224 low-order symbols will be described. The representative symbol frequency estimation unitcalculates an estimated value S of the sum of the frequencies of occurrence of the low-order symbols by the following equation (2) using a frequency of occurrence F(i) of a symbol indicated in an entry in the second part tablethat is identified by an index i.
432 In the calculation according to the equation (2), the 224 entries included in the second part tableare divided into entries for every 16 indexes from the head, and 14 ranges (hereinafter referred to as index ranges) are set. A k-th index range from the head among the 14 index ranges is referred to as a k-th index range. Note that k is any value from 0 to 13.
352 352 352 The representative symbol frequency estimation unitcalculates an estimated value of a frequency of occurrence for each of the 0-th to 13-th index ranges. Specifically, the representative symbol frequency estimation unitmultiplies an average value of a frequency of occurrence F (k×16) indicated in the first entry of the k-th index range and a frequency of occurrence F (k×16+15) indicated in the last entry of the k-th index range by the number of symbols in one index range (here, 16), thereby calculating the estimated value of the frequency of occurrence corresponding to the k-th index range. Then, the representative symbol frequency estimation unitcalculates the sum of the estimated values of the frequencies of occurrence that correspond to the 0-th to 13-th index ranges, respectively, thereby obtaining an estimated value S of the sum of the frequencies of occurrence of the low-order symbols.
432 In the calculation according to the equation (2), the number of times of addition required to calculate the estimated value S is reduced by assuming that the frequencies of occurrence corresponding to the respective indexes vary linearly from the head entry to the last entry in each index range. Note that the method of calculating the estimated value S is an example, and other methods may be used. For example, the number of index ranges set by dividing the 224 entries included in the second part tablemay be appropriately changed in consideration of a calculation amount required to calculate the estimated value S and accuracy of the calculated estimated value S to be estimated.
352 324 The representative symbol frequency estimation unitsends the estimated value S of the sum of the frequencies of occurrence of the low-order symbols to the Huffman tree generation unitas a frequency of occurrence of the representative symbol.
324 431 324 431 11 431 324 325 324 3 FIG. 3 FIG. The Huffman tree generation unitgenerates a Huffman tree by using the frequency of occurrence of each of the H non-representative symbols in the first part tableand the frequency of occurrence of the representative symbol. Specifically, the Huffman tree generation unitarranges, as leaf nodes, symbols each having a frequency of occurrence larger than 0 among the H non-representative symbols in the first part tableand the representative symbol and then performs a Huffman tree generation process. This Huffman tree generation process is similar to the Huffman tree generation process described above with reference to. Specifically, this Huffman tree generation process is the Huffman tree generation process described above with reference toin which step Sis replaced with a step of adding, as leaf nodes on the Huffman tree, symbols each having a frequency of occurrence larger than 0 among the H non-representative symbols in the first part tableand the representative symbol. The Huffman tree generation unitsends the generated Huffman tree to the code length determination unit. Note that in a case where there are M representative symbols, the Huffman tree generation unitgenerates the Huffman tree, based on frequencies of occurrence of symbols that are larger than 0 among the H non-representative symbols and the M representative symbols.
16 FIG. 60 324 60 600 610 611 621 631 641 642 643 644 645 shows an example of the Huffman treegenerated by the Huffman tree generation unit. The Huffman treeincludes a root node, intermediate nodes,,, and, first-type leaf nodes,,, and, and a second-type leaf node.
641 642 643 644 431 Each of the first-type leaf nodes,,, andcorresponds to a symbol having a frequency of occurrence larger than 0 among the symbols (non-representative symbols) included in the first part table. The number of first-type leaf nodes is at most H.
645 The second-type leaf nodecorresponds to the representative symbol whose frequency of occurrence is larger than 0. The number of second-type leaf nodes is at most M (here, M=1).
324 324 32 324 32 324 32 Thus, the number of the symbols on which the Huffman tree generation is performed (i.e., the number of leaf nodes of the Huffman tree to be generated) by the Huffman tree generation unitis at most (H+M). In contrast, as described above, the number of symbols on which the Huffman tree generation is performed by the Huffman tree generation unitA of the code table generation unitA according to the comparative example is at most N. For example, in a case where N is 256, H is 32, and M is 1, the number of the symbols on which the Huffman tree generation is performed by the Huffman tree generation unitis significantly reduced from at most 256 to at most 33 (=32+1) in the code table generation unitof the present embodiment as compared to the Huffman tree generation unitA of the comparative example. Accordingly, in the code table generation unit, it is possible to reduce a time period required for the code table generation process or to reduce a circuit scale (for example, the number of gates) for completing the code table generation within a specific time period.
432 431 431 432 However, as a result of all the symbols included in the second part tablebeing treated as one representative symbol, there is a concern that coding efficiency of dynamic Huffman coding is lowered. In order to address the concern, the H high-order symbols (non-representative symbols) that are included in the first part tableand account for most of the frequencies of occurrence are treated as H symbols (that is, H leaf nodes) also on the Huffman tree. Accordingly, accuracy of the dynamic Huffman coding is secured. In other words, the number H of symbols included in the first part tableis reduced within a range in which the accuracy of the dynamic Huffman coding can be secured. In addition, the L low-order symbols included in the second part tableoften have small frequencies of occurrence. Thus, even if a non-optimal code lengths are set for the low-order symbols, influence on the coding efficiency is small.
32 Therefore, the code table generation unitof the present embodiment can reduce a time period required for the code table generation process or can reduce a circuit scale for completing the code table generation within a specific time, while keeping decrease in the coding efficiency within a practically acceptable range.
353 432 432 432 432 353 353 354 The merge symbol count unitcounts the number C of symbols in the second part tablethat have frequencies of occurrence larger than 0, respectively. C is an integer equal to or larger than 0. If the second part tableincludes no symbol whose frequency of occurrence is larger than 0, C is 0. A symbol in the second part tablewhose frequency of occurrence is larger than 0 is referred to as a merge symbol. The merge symbol is a symbol to which a variable length code needs to be assigned among the symbols included in the second part table. The number C of merge symbols counted by the merge symbol count unitis the number of merge symbols that are represented by the representative symbol. The merge symbol count unitsends the number C of merge symbols to the merge symbol additional code length determination unit.
354 432 354 327 354 326 The merge symbol additional code length determination unitdetermines, for each merge symbol, a code length to be added to the code length of the representative symbol (hereinafter, referred to as an additional code length) by using a subtree that is based on the second part tableand the number C of merge symbols. The merge symbol additional code length determination unitsends the additional code length for each of the C merge symbols to the representative symbol expansion unit. In addition, the merge symbol additional code length determination unitalso sends the maximum value of the additional code lengths of the C merge symbols to the maximum code length restriction unit.
60 A code length of each merge symbol is obtained by adding its additional code length to the code length of the representative symbol. This is equivalent to determining a structure of a subtree having leaf nodes to which the respective merge symbols are assigned, and determining a position on the Huffman treefor the root node of the subtree by the Huffman tree generation process.
17 FIG. 65 354 65 645 60 shows an example of a subtreeof merge symbols that is used by the merge symbol additional code length determination unit. The subtreeis a binary tree whose root node is the leaf nodeof the Huffman tree. Here, a case where the number C of merge symbols is five will be exemplified.
60 600 611 621 630 631 640 641 642 643 644 645 60 16 FIG. The Huffman treeincludes the root node, the intermediate nodes,,, and, the first-type leaf nodes,,,, and, and the second-type leaf node, similarly to the Huffman treedescribed above with reference to.
641 642 643 644 431 Each of the first-type leaf nodes,,, andcorresponds to a symbol having a frequency of occurrence larger than 0 among the symbols (non-representative symbols) included in the first part table.
645 645 65 645 The second-type leaf nodecorresponds to a representative symbol whose frequency of occurrence larger than 0. The second-type leaf nodeis the root node of the subtree. The frequency of occurrence associated with the second-type leaf nodeis the estimated value S of the sum of the frequencies of occurrence of the merge symbols.
65 645 60 650 651 660 661 662 663 670 671 661 662 663 670 671 The subtreeincludes the leaf nodeof the Huffman treeas its root node and includes intermediate nodes,, andand leaf nodes,,,, and. The leaf nodes,,,, andcorrespond to the five merge symbols, respectively.
661 662 663 670 671 645 65 645 Thus, a code length of each merge symbol is determined by adding, as an additional code length, a depth (the number of edges) of each of the leaf nodes,,,, andstarting from the root node (i.e., the leaf node) in the subtreeto a code length of the representative symbol corresponding to the leaf node.
65 From the viewpoint of reduction in a processing amount, it is desirable that the subtreeincluding C leaf nodes that correspond to C merge symbols, respectively, has a structure in which the additional code length for each merge symbol can be easily obtained by using the number C of merge symbols.
354 65 65 354 65 354 65 Thus, the merge symbol additional code length determination unitadopts, for example, a balanced binary tree as the structure of the subtree. The balanced binary tree is a binary tree in which, for all the leaf nodes, a difference in a depth between the leaf nodes is at most one. Note that a structure other than the balanced binary tree may be used as the subtree. The merge symbol additional code length determination unitmay determine the structure of the subtree, based on the frequencies of occurrence of the C merge symbols. In addition, the merge symbol additional code length determination unitmay also determine the structure of the subtreeby selecting from tree structures (for example, templates of tree structure) prepared in advance.
65 354 When using the balanced binary tree as the structure of the subtree, the merge symbol additional code length determination unitdetermines the additional code length for each merge symbol by the following procedure (A1) and (A2).
X (A1) Calculate a minimum integer X that satisfies 2≥C.
X X 432 (A2) Set additional code lengths of (2−C) merge symbols in ascending order of the indexes on the second part table(that is, in descending order of the frequencies of occurrence) among the C merge symbols to (X−1) bits, and set additional code lengths of the remaining (2C−2) merge symbols to X bits. X is the maximum value of the additional code lengths of the C merge symbols.
17 FIG. 354 354 432 661 662 663 354 670 671 3 3 3 In the example shown in, C=5. Thus, in a case where X is calculated according to the above-described procedure, the merge symbol additional code length determination unitcalculates X=3 since 2>5. Then, the merge symbol additional code length determination unitsets the addition code lengths of three (=2−5) merge symbols in ascending order of the indexes on the second part tableamong the five merge symbols (corresponding to the leaf nodes,, and) to 2 (=3−1) bits. In addition, the merge symbol additional code length determination unitsets additional code lengths of the remaining two (=2×5−2) merge symbols (corresponding to the leaf nodesand) to 3 bits. The maximum value of the additional code lengths of the five merge symbols is three.
Here, a function f(K(r)) that derives the maximum value of additional code lengths corresponding to a symbol set will be considered in a case where r is a representative symbol that represents the symbol set and K(r) is the number of symbols having frequencies of occurrence larger than 0 (i.e., the number of merge symbols) among symbols included in the symbol set represented by the representative symbol r. That is, f is a function of K(r). The following equation (3) defines f(K(r)).
354 The merge symbol additional code length determination unitmay calculate the maximum value of the additional code lengths of the C merge symbols by this function f(K(r)) in which C is substituted for K(r) (i.e., K(r)=C).
12 FIG. The description returns to.
60 324 325 431 325 By using the Huffman treegenerated by the Huffman tree generation unit, the code length determination unitdetermines a code length of each of symbols having frequencies of occurrence larger than 0 among the H non-representative symbols (high-order symbols) included in the first part table, and a code length of the representative symbol. Hereinafter, in order to make the description easy to understand, an example in which the frequencies of occurrence of all the H non-representative symbols are larger than 0 will be described. Note that in practice, the H non-representative symbols may include a symbol whose frequency of occurrence is 0. In this case, the code length determination unitdoes not calculate a code length of a symbol whose frequency of occurrence is 0 among the H non-representative symbols.
325 600 60 640 643 645 17 FIG. Specifically, the code length determination unitdetermines the number of edges passing from a leaf node corresponding to each of the symbols, which include the H non-representative symbols and the representative symbol, to the root nodeby tracing edges in the Huffman tree, as the code length of the symbol. In the example shown in, for example, the code length of the non-representative symbol corresponding to the leaf nodeis 1 bit. For example, the code length of the non-representative symbol corresponding to the leaf nodeis 4 bits. In addition, for example, the code length of the representative symbol corresponding to the leaf nodeis 3 bits.
325 326 The code length determination unitsends the determined H combinations of the non-representative symbol and the code length, and the combination of the representative symbol and the code length, to the maximum code length restriction unit. Hereinafter, a combination of a non-representative symbol and a code length is also referred to as a <non-representative symbol, code length> pair. A combination of a representative symbol and a code length is also referred to as a <representative symbol, code length> pair.
326 354 325 354 325 326 The maximum code length restriction unitexecutes a process of restricting a code length corresponding to a symbol to the maximum code length (maximum code length restriction process) by using the maximum value of the additional code lengths sent from the merge symbol additional code length determination unit, and the H <non-representative symbol, code length> pairs and the <representative symbol, code length> pair that are sent from the code length determination unit. By using the maximum value of the additional code lengths sent from the merge symbol additional code length determination unitand the <representative symbol, code length> pair sent from the code length determination unit, the maximum code length restriction unitobtains a combination of the representative symbol, the code length, and the maximum value of the additional code lengths. Hereinafter, the combination of the representative symbol, the code length, and the maximum value of the additional code lengths is also referred to as a <representative symbol, code length, maximum value of additional code lengths> pair.
18 FIG. 326 326 371 372 373 374 375 376 377 is a block diagram showing an example of a configuration of the maximum code length restriction unit. The maximum code length restriction unitincludes, for example, a code length sorting unit, a code length clipping unit, a representative symbol swapping unit, a violation symbol count calculation unit, a termination determination unit, a code length change unit, and a violation symbol count decrement unit.
371 First, the code length sorting unitsorts the H <non-representative symbol, code length> pairs and the <representative symbol, code length, maximum value of additional code lengths> pair in descending order of the code lengths.
372 372 372 376 Next, the code length clipping unitchanges (i.e., clips) a code length longer than the maximum code length included in the H <non-representative symbol, code length> pairs, to the maximum code length. Specifically, the code length clipping unitidentifies all combinations each including a code length longer than the maximum code length, among the H <non-representative symbol, code length> pairs. The code length clipping unitchanges the code length included in each of the identified combinations, to the maximum code length. This change causes the H <non-representative symbol, code length> pairs to include non-representative symbols corresponding to illegally short code lengths. In other words, variable length codes assigned based on the H <non-representative symbol, code length> pairs are invalid Huffman codes. Such illegally short code lengths are corrected by the code length change unitto be described later.
372 Note that in a case where the H <non-representative symbol, code length> pairs do not include any combination including a code length longer than the maximum code length, the code length clipping unitdoes not perform a process on the H <non-representative symbol, code length> pairs and the <representative symbol, code length, maximum value of additional code lengths> pair.
373 373 Next, when the code length of the representative symbol violates the restriction on the maximum code length, the representative symbol swapping unitperforms a process of swapping (exchanging) the code lengths between the representative symbol and a non-representative symbol that satisfies a specific condition (swap process). The representative symbol swapping unitdetermines whether or not the code length of the representative symbol violates the restriction on the maximum code length by using the following inequality (4).
Code lengths of the merge symbols represented by the representative symbol are at most a length obtained by adding the maximum value of the additional code lengths to the code length of the representative symbol. Therefore, when the code length of the representative symbol does not satisfy this inequality, it is guaranteed that the code lengths of all the merge symbols represented by the representative symbol are shorter than or equal to the maximum code length.
373 In contrast, when the code length of the representative symbol satisfies this inequality, the representative symbol swapping unitexecutes a process of swapping the code lengths between the representative symbol and the non-representative symbol that satisfies the specific condition. The non-representative symbol that satisfies the specific condition is a non-representative symbol that satisfies the following inequality (5) and has the longest code length.
373 In other words, the representative symbol swapping unitexecutes the swap process when the code length of the representative symbol is longer than a representative symbol maximum code length.
The representative symbol maximum code length indicates an upper limit of the code length of the representative symbol. The representative symbol maximum code length is set such that the code lengths of all the merge symbols represented by the representative symbol are restricted to the maximum code length. Therefore, the representative symbol maximum code length is obtained by subtracting the maximum value of the additional code lengths from the maximum code length, as shown below in equation (6). Note that the number of the merge symbols represented by the representative symbol is denoted by C.
When the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length, it is guaranteed that the code lengths of all the merge symbols represented by the representative symbol are shorter than or equal to the maximum code length.
373 In contrast, when the code length of the representative symbol is longer than the representative symbol maximum code length, the code length of at least one of the merge symbols represented by the representative symbol is longer than the maximum code length. Therefore, the representative symbol swapping unitswaps the code length of the representative symbol for the longest code length among the code lengths of the non-representative symbols that are shorter than the representative symbol maximum code length.
373 373 373 More specifically, when the code length of the representative symbol is longer than the representative symbol maximum code length, the representative symbol swapping unitidentifies non-representative symbols whose code lengths are shorter than or equal to the representative symbol maximum code length, by using the H <non-representative symbol, code length> pairs. The representative symbol swapping unitfurther identifies a non-representative symbol having the longest code length (hereinafter referred to as a swap target non-representative symbol) among the non-representative symbols whose code lengths are shorter than or equal to the representative symbol maximum code length. The representative symbol swapping unitswaps the code lengths between the <representative symbol, code length> pair (i.e., <representative symbol, code length, maximum value of additional code lengths> pair) and the <swap target non-representative symbol, code length> pair.
373 Note that when the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length, the representative symbol swapping unitdoes not execute a process for the H <non-representative symbol, code length> pairs and the <representative symbol, code length, maximum value of additional code lengths> pair.
374 374 Next, the violation symbol count calculation unitcalculates the number I of violation symbols, based on the H <non-representative symbol, code length> pairs and the <representative symbol, code length, maximum value of additional code lengths> pair. The number I of violation symbols indicates the number of symbols to which variable length codes cannot be assigned. When the number I of violation symbols is larger than 0, Huffman codes determined based on the H <non-representative symbol, code length> pairs and the <representative symbol, code length, maximum value of additional code lengths> pair are invalid Huffman codes. The violation symbol count calculation unitcalculates the number I of violation symbols by the following equation (7).
Here, symbol set AllSymbol includes the H non-representative symbols and the representative symbol. lmax is the maximum code length.
374 Note that the violation symbol count calculation unitmay calculate the number I of violation symbols by the following equation (8).
372 The equation (8) represents an equation of deriving the number I of violation symbols that includes the clipping of the code lengths of the non-representative symbols executed by the code length clipping unit. clip(l(s)) is a function of clipping code length (l(s)) of symbol s to maximum code length lmax. When code length l(s) is longer than maximum code length lmax, clip(l(s)) is maximum code length lmax (i.e., clip(l(s))=lmax). When code length l(s) is shorter than or equal to maximum code length lmax, clip(l(s)) is code length l(s) (i.e., clip(l(s))=l(s)).
375 326 375 375 The termination determination unitdetermines whether or not to terminate the maximum code length restriction process in the maximum code length restriction unit, based on the number I of violation symbols. Specifically, when the number I of violation symbols is 0, the termination determination unitdetermines that the maximum code length restriction process is terminated. When the number I of violation symbols is larger than 0, the termination determination unitdetermines that the maximum code length restriction process is continued.
375 376 377 When the termination determination unithas determined that the maximum code length restriction process is continued, the following processes are executed by the code length change unitand the violation symbol count decrement unit.
376 376 The code length change unitchanges the code lengths of the non-representative symbols such that the number I of violation symbols decreases by using the H <non-representative symbol, code length> pairs. Specifically, the code length change unitchanges the code lengths of the non-representative symbols according to the following steps (B1) to (B4).
372 (B1) Select a violation symbol s by using the H <non-representative symbol, code length> pairs. The violation symbol s is a non-representative symbol that corresponds to a code length equal to the maximum code length. The violation symbol s is, for example, a symbol whose code length was longer than the maximum code length but has been changed to the maximum code length by the code length clipping unit.
(B2) Select a non-representative symbol s′ whose code length is shorter than the maximum code length and is longer by using the H <non-representative symbol, code length> pairs.
(B3) Change the code length l(s) of the violation symbol s to a length that is obtained by adding one to the code length l(s′) of the non-representative symbol s′.
(B4) Change the code length l(s′) of the non-representative symbol s′ by incrementing the length by one.
Note that the steps (B3) and (B4) may be replaced with other steps that achieves similar results with respect to changing the code lengths. For example, the steps (B3) and (B4) may be replaced with the following steps (B3′) and (B4′).
(B3′) Change the code length l(s′) of the non-representative symbol s′ by incrementing the length by one.
(B4′) Change the code length l(s) of the violation symbol s to a length that is equal to the changed code length l(s′) of the non-representative symbol s′.
376 377 After the code length l(s) of the violation symbol s and the code length l(s′) of the non-representative symbol s′ are changed by the code length change unit, the violation symbol count decrement unitdecrements the number I of violation symbols by one thereby updating the number I of violation symbols.
375 326 376 377 326 327 The termination determination unitdetermines whether or not to terminate the maximum code length restriction process in the maximum code length restriction unit, based on the updated number I of violation symbols. Therefore, the processes of the code length change unitand the violation symbol count decrement unitare repeated until the number I of violation symbols becomes 0. When the number I of violation symbols has reached 0, the H <non-representative symbol, code length> pairs and the <representative symbol, code length> pair that have been changed to code lengths from which valid Huffman codes are generated, are obtained. The maximum code length restriction unitsends the obtained H <non-representative symbol, code length> pairs and <representative symbol, code length> pair to the representative symbol expansion unit.
326 19 FIG. 23 FIG. A specific example in which code lengths are restricted to the maximum code length in the maximum code length restriction unitwill be described with reference toto. Here, it is assumed that the maximum code length is 4 bits.
19 FIG. 19 FIG. shows an example of non-representative symbols and a representative symbol determined based on frequencies of occurrence. In the example shown in, the bottom four symbols “h”, “i”, “j”, and “k” among symbols arranged in descending order of frequencies of occurrence are represented by a representative symbol R. In other words, the four symbols “h”, “i”, “j”, and “k” are merge symbols. The sum of the frequencies of occurrence of the four merge symbols “h”, “i”, “j”, and “k”, i.e., four is set as a frequency of occurrence of the representative symbol “R”. The remaining seven symbols “a”, “b”, “c”, “d”, “e”, “f”, and “g” (hereinafter referred to as non-representative symbols “a” to “g”) are non-representative symbols.
20 FIG. 19 FIG. 3 FIG. 70 70 70 11 shows an example of a Huffman treecorresponding to the non-representative symbols and the representative symbol shown in. The Huffman treeis generated based on the frequency of occurrence of each of the seven non-representative symbols “a” to “g” and the frequency of occurrence of the representative symbol “R”. Specifically, the Huffman treeis generated by, for example, the Huffman tree generation process described above with reference toin which step Sis replaced with a step of adding, as leaf nodes, the seven non-representative symbols “a” to “g” and the representative symbol “R”.
326 326 326 Note that in the description on the process in the maximum code length restriction unit, an example in which a structure of a Huffman tree is shown and symbols are assigned to leaf nodes is illustrated in order to make change of code lengths easy to understand. Then, even when an example in which a structure of a Huffman tree is shown and symbols are assigned to leaf nodes is illustrated, the maximum code length restriction unitmay not actually manage the structure of the Huffman tree and the relationship between each leaf node and each symbol. The maximum code length restriction unitmanages, for example, the <representative symbol, code length> pair and a list of the <non-representative symbol, code length> pairs, thereby performing the selection of symbols and the change of code lengths.
70 712 722 742 743 753 765 766 741 In the generated Huffman tree, the non-representative symbol “a” is assigned to a leaf node. The non-representative symbol “b” is assigned to a leaf node. The non-representative symbols “c” and “d” are assigned to leaf nodesand, respectively. The non-representative symbol “e” is assigned to a leaf node. The non-representative symbols “f” and “g” are assigned to leaf nodesand, respectively. The representative symbol “R” is assigned to a leaf node.
A code length of the non-representative symbol “a” is 1 bit. A code length of the non-representative symbol “b” is 2 bits. A code length of the representative symbol “R” and code lengths of the two non-representative symbols “c” and “d” are 4 bits. A code length of the non-representative symbol “e” is 5 bits. Code lengths of the two non-representative symbols “f” and “g” are 6 bits.
75 75 761 762 763 764 Note that the edges and nodes represented by dotted lines indicate a subtreecorresponding to the merge symbols “h”, “i”, “j”, and “k” (hereinafter referred to as merge symbols “h” to “k”). The subtreeis a subtree having the representative symbol “R” as its root node and having each of the merge symbols “h” to “k” as a leaf node. The merge symbols “h” to “k” are assigned to leaf nodes,,, and, respectively. A code length of each of the merge symbols “h” to “k” is 6 bits. Therefore, for these merge symbols “h” to “k”, the maximum value of the code lengths which are to be added to the code length of the representative symbol “R” (i.e., the maximum value of additional code lengths) is two (=6−4).
70 372 In the Huffman Tree, the code lengths of the non-representative symbols “e”, “f”, and “g” are longer than the maximum code length. Accordingly, the code length clipping unitchanges the code lengths of the non-representative symbols “e”, “f”, and “g” to the maximum code length.
70 373 373 Furthermore, in the Huffman tree, a length obtained by adding the maximum value of the additional code lengths to the code length of the representative symbol “R” is longer than the maximum code length (that is, 4+2 >4). Accordingly, the representative symbol swapping unitselects the non-representative symbol “b” having the longest code length among the non-representative symbols “a” and “b” each having a code length shorter than or equal to a representative symbol maximum code length. The representative symbol maximum code length is obtained by subtracting the maximum value of the additional code lengths from the maximum code length (=4−2=2). The representative symbol swapping unitexchanges the code lengths between the representative symbol “R” and the selected non-representative symbol “b”.
21 FIG. 20 FIG. 70 70 372 373 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the code length clipping unithas changed the code lengths of the non-representative symbols “e”, “f”, and “g” to the maximum code length and the representative symbol swapping unithas exchanged the code lengths between the representative symbol “R” and the non-representative symbol “b” as described above.
70 744 745 746 In other words, in the modified Huffman tree, the code lengths of the non-representative symbols “e”, “f”, and “g” (leaf nodes,, and) are set to the maximum code length (=4). The code length of the representative symbol “R” is set to the original code length of the non-representative symbol “b” (=2). The code length of the non-representative symbol “b” is set to the original code length of the representative symbol “R” (=4). As a result, the merge symbols “h” to “k” represented by the representative symbol “R” each have a code length shorter than or equal to the maximum code length, thus satisfying the restriction on the maximum code length.
745 746 701 745 746 Note that the leaf nodesandare nodes that cannot be traced from the root node. Consequently, the non-representative symbols “f” and “g” assigned to the leaf nodesand, respectively, are violation symbols.
374 In this case, the violation symbol count calculation unitcalculates the number I of violation symbols according to the above-described equation (7) (or equation (8)) as described below.
70 376 21 FIG. Therefore, in the Huffman treeshown in, the number I of violation symbols is two. Since the calculated number I of violation symbols is larger than 0, the code length change unitchanges the code lengths of the non-representative symbols until the number I of violation symbols becomes 0.
376 376 376 Specifically, the code length change unitselects, for example, the violation symbol “f”. Then, the code length change unitselects the non-representative symbol “a” whose code length is shorter than the maximum code length. Note that when there are multiple non-representative symbols whose code lengths are shorter than the maximum code length, the code length change unitselects the first non-representative symbol from the multiple non-representative symbols in the order of longer code lengths.
376 376 The code length change unitchanges the code length of the violation symbol “f” to a length obtained by adding one to the code length of the non-representative symbol “a”. Then, the code length modification unitincrements the code length of the non-representative symbol “a” by one. As a result, the code length of the violation symbol “f” and the code length of the non-representative symbol “a” become 2 bits.
377 375 Next, the violation symbol count decrement unitdecrements the number I of violation symbols by one, thereby updating the number I of violation symbols with one. Since the updated number I of violation symbols is larger than 0, the termination determination unitdecides to further change the code lengths of the non-representative symbols.
22 FIG. 21 FIG. 70 70 376 70 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the code length change unithas changed the code lengths of the non-representative symbol “a” and the violation symbol “f as described above. In other words, in the modified Huffman Tree, the code length of the non-representative symbol “a” and the code length of the violation symbol “f” are set to 2 bits. The non-representative symbol “g” is still a violation symbol.
376 375 376 376 376 376 The code length change unitselects the violation symbol “g” in response to the decision of further changing the code lengths of the non-representative symbols by the termination determination unit. Then, the code length change unitselects the non-representative symbol “f” having the longest code length among the non-representative symbols “a” and “f” whose code lengths are shorter than the maximum code length. In this case, since both the code lengths of the non-representative symbols “a” and “f” are 2 bits, the code length change unitmay select the non-representative symbol “a”. When there are multiple non-representative symbols each having the longest code length among non-representative symbols whose code lengths are shorter than the maximum code length, the code length change unitselects one of the multiple non-representative symbols under specific rules. In the following descriptions, it is assumed that the code length change unithas selected the non-representative symbol “f”.
376 376 The code length change unitchanges the code length of the selected violation symbol “g” to a length obtained by adding one to the code length of the selected non-representative symbol “f”. Then, the code length change unitincrements the code length of the non-representative symbol “f” by one. As a result, the code length of the violation symbol “g” and the code length of the non-representative symbol “f” become 3 bits.
377 375 Next, the violation symbol count decrement unitdecrements the number I of violation symbols by one, thereby updating the number I of violation symbols with 0. Since the updated number I of violation symbols is 0, the termination determination unitdecides to terminate the change of the code lengths of the non-representative symbols.
23 FIG. 22 FIG. 70 70 376 70 70 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the code length change unithas changed the code lengths of the non-representative symbol “f” and the violation symbol “g” as described above. In other words, in the modified Huffman tree, the code length of the non-representative symbol “f” and the code length of the violation symbol “g” are set to 3 bits. In addition, this Huffman treehas no violation symbols.
70 70 70 70 This Huffman treeis a valid Huffman tree and every intermediate node of the Huffman treehas two child nodes. Therefore, Huffman codes determined based on the Huffman treeare valid Huffman codes that satisfy Kraft's inequality and have no redundant code assignment (that is, K=1). In other words, the Huffman codes determined based on the Huffman treeare perfect codes.
24 FIG. 25 FIG. An example of the procedure of the maximum code length restriction process will be described with reference toand.
24 FIG. 326 326 325 is a flowchart showing an example of the procedure of the maximum code length restriction process executed in the maximum code length restriction unit. The maximum code length restriction unitexecutes the maximum code length restriction process for the H <non-representative symbol, code length> pairs and the <representative symbol, code length> pair sent from the code length determination unit.
326 301 326 302 First, the maximum code length restriction unitchanges the code lengths of all non-representative symbols that are longer than the maximum code length, to the maximum code length (step S). Next, the maximum code length restriction unitdetermines whether or not the code length of the representative symbol is longer than the representative symbol maximum code length (step S). The representative symbol maximum code length is a length obtained by subtracting the maximum value of the additional code lengths from the maximum code length.
302 326 303 326 304 305 When the code length of the representative symbol is longer than the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitselects a non-representative symbol having the longest code length among the non-representative symbols each having a code length shorter than or equal to the representative symbol maximum code length (step S). The maximum code length restriction unitswaps the code lengths between the selected non-representative symbol and the representative symbol (step S) and proceeds to step S.
302 326 305 When the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length (No in step S), the process by the maximum code length restriction unitproceeds to step S.
326 305 Next, the maximum code length restriction unitcalculates the number I of violation symbols according to the above-described equation (7) (or equation (8)) (step S).
326 306 The maximum code length restriction unitdetermines whether or not the calculated number I of violation symbols is 0 (step S). The number I of violation symbols larger than 0 indicates that Huffman codes assigned based on the current code length of each symbol are invalid Huffman codes.
306 326 307 25 FIG. When the number I of violation symbols is larger than 0 (No in step S), the maximum code length restriction unitexecutes a first code length change process (step S). The first code length change process is a process of changing a code length of one violation symbol to a code length shorter than or equal to the maximum code length. A specific procedure of the first code length change process will be described later with reference to a flowchart of.
326 308 326 326 306 326 307 308 Next, the maximum code length restriction unitdecrements the number I of violation symbols by one (step S). That is, the maximum code length restriction unitsubtracts one from the number I of violation symbols. Then, the process by the maximum code length restriction unitreturns to step Sof determining whether or not the number I of violation symbols is 0. Therefore, the maximum code length restriction unitrepeats step Sof executing the first code length change process and step Sof decrementing the number I of violation symbols, until the number I of violation symbols becomes 0.
306 326 When the number I of violation symbols is 0 (Yes in step S), the maximum code length restriction unitends the maximum code length restriction process.
326 With the above maximum code length restriction process, the maximum code length restriction unitcan restrict the code lengths of the non-representative symbols to code lengths that are shorter than or equal to the maximum code length, and restrict the code length of the representative symbol to a code length that is shorter than or equal to the representative symbol maximum code length. By restricting the code length of the representative symbol to the code length shorter than or equal to the representative symbol maximum code length, the code lengths of the symbols represented by the representative symbol (i.e., the code lengths of the merge symbols) can be restricted to code lengths that are shorter than or equal to the maximum code length.
25 FIG. 24 FIG. 326 307 326 is a flowchart showing an example of the procedure of the first code length restriction process executed in the maximum code length restriction unit. The first code length change process corresponds to step Sof the maximum code length restriction process described above with reference to. The maximum code length restriction unitexecutes the first code length change process when the non-representative symbols include one or more violation symbols.
326 41 326 42 326 326 43 326 44 326 First, the maximum code length restriction unitselects a non-representative symbol s′ having the longest code length among the non-representative symbols each having a code length shorter than the maximum code length (step S). The maximum code length restriction unitselects one violation symbol from the one or more violation symbols (step S). Specifically, the maximum code length restriction unitselects one non-representative symbol having a code length equal to the maximum code length, as the violation symbol. The maximum code length restriction unitchanges the code length of the selected violation symbol to a length obtained by adding one to the code length l(s′) of the non-representative symbol s′ (step S). Then, the maximum code length restriction unitincrements the code length l(s′) of the non-representative symbol s′ by one (step S) and ends the first code length change process. That is, the maximum code length restriction unitadds one to the code length l(s′) of the non-representative symbol s′ and ends the first code length change process.
326 326 326 326 The above-described first code length change process corresponds to the operation in which the maximum code length restriction unitchanges the leaf node corresponding to the non-representative symbol s′ to an intermediate node, and assigns the non-representative symbol s′ and the violation symbol to two leaf nodes, respectively, which have the intermediate node as their parent. In this operation, no leaf node to which no symbol is assigned is generated. That is, no redundant leaf node is generated. For this reason, the maximum code length restriction unitcan guarantee that the Huffman codes assigned based on the code length of each symbol are perfect codes when the maximum code length restriction process including the first code length change process has been completed. Therefore, the maximum code length restriction unitcan reduce the processing time and the processing resources as compared to, for example, the maximum code length restriction unitA of the comparative example, which additionally executes a process of merging a binary tree having a redundant leaf node. Furthermore, since the maximum code length restriction process is executed on the representative symbol that represents the merge symbols, the number of symbols handled in the maximum code length restriction process can be significantly reduced. Therefore, the processing time and the processing resources can also be reduced by the reduction in the number of symbols.
12 FIG. The description returns to.
327 326 327 354 327 The representative symbol expansion unitreceives the H <non-representative symbol, code length> pairs and the <representative symbol, code length> pair for which the maximum code length restriction process has been completed, from the maximum code length restriction unit. The representative symbol expansion unitreceives the C <merge symbol, additional code length> pairs from the merge symbol additional code length determination unit. The representative symbol expansion unitdetermines the code length of each of the C merge symbols by using the received <representative symbol, code length> pair and the C <merge symbol, additional code length> pairs.
327 327 327 328 326 328 The representative symbol expansion unitdetermines the code length of each of the C merge symbols by using the code length of the representative symbol and the C additional code lengths that correspond to the C merge symbols, respectively. Specifically, for each of the C merge symbols, the representative symbol expansion unitadds an additional code length corresponding to a merge symbol to the code length of the representative symbol, thereby determining the code length of the merge symbol. The representative symbol expansion unitsends a list including the H <non-representative symbol, code length> pairs and C <merge symbol, code length> pairs, to a code assignment unit. Note that the H <non-representative symbol, code length> pairs received from the maximum code length restriction unitare sent to the code assignment unitas they are.
328 327 The code assignment unitgenerates a code table according to, for example, the canonical Huffman coding by using the list including the H <non-representative symbol, code length> pairs and the C <merge symbol, code length> pairs sent by the representative symbol expansion unit. The canonical Huffman coding is capable of determining a code bit string (variable length code) to be assigned to each of symbols by using only code lengths of the symbols. In the canonical Huffman coding, a code bit string is assigned to each symbol according to the following rules (C1) and (C2) after sorting the list by using the code lengths as the first key and using the symbol numbers as the second key.
(C1) A code bit string to be assigned to a symbol having a short code length precedes, in lexicographic order, a code bit string to be assigned to a symbol having a long code length.
(C2) For any two symbols whose code lengths are equal, a code bit string to be assigned to one of the two symbols preceding in symbol order precedes, in the lexicographic order, a code bit string to be assigned to the other of the two symbols succeeding in the symbol order.
For example, in order to determine the lexicographic order of the code bit strings, the lexicographic order of bit values “0” and “1” is defined as an order of “0” and “1”. In this case, each bit value included in a code bit string and each bit value included in another code bit string are compared in order from the higher bit to determine the order (bit order) of the corresponding bit values, thereby determining the lexicographic order of the code bit strings.
More specifically, for example, “1′b0”, “2′b10”, “3′b110”, and “3′b111” are four code bit strings arranged in the lexicographic order. Note that a data string including at least one bit value of 0 or 1 subsequent to “X′b” indicates a bit data string of X bits. Thus, in these four code bit strings, the 1-bit code bit string “1′b0” precedes the 2-bit code bit string “2′b10” in the bit order of the most significant bits. The 2-bit code bit string “2′b10” precedes the 3-bit code bit string “3′b110” in the bit order of the next most significant bits. The 3-bit code bit string “3′b110” precedes the 3-bit code bit string “3′b111” in the bit order of the least significant bits.
26 FIG. 328 shows an example of a pseudo-program for the code assignment unitto assign a code bit string to a symbol. In the pseudo-program, a variable “code” is used for calculating a code bit string to be assigned to a symbol.
328 First, the code assignment unitsets the variable “code” to 0 (=1′b0). When code bit strings are to be assigned to symbols, respectively, the number of bits of the variable “code” initially depends on a minimum code length of code lengths that are associated with the symbols, respectively.
328 Next, in one while loop, the code assignment unitdetermines a code bit string to be assigned to one symbol in ascending code length order, and in specific symbol order in the case of the same code length.
328 328 328 For example, in a first while loop, the code assignment unitselects, as a target to which a code bit string is assigned, a symbol that has the shortest code length and precedes, in specific symbol order, another symbol having the same code length, if any. Then, the code assignment unitassigns the variable “code” (=1′b0) as a code bit string of the selected symbol. As described above, the number of bits of the variable “code” assigned first depends on the minimum code length. In this example, it is assumed that the minimum code length is one. Then, the code assignment unitperforms a shift operation of shifting a value that is obtained by adding one to the variable “code”, to the left by the number of bits obtained by subtracting the code length of the current symbol from the code length of the next symbol, and sets the value obtained by the shift operation as the variable “code”. For example, in a case where the number of bits obtained by subtracting the code length of the current symbol from the code length of the next symbol is one, a value 2′b10 obtained by shifting 1′b1 that is a value obtained by adding one to the variable “code” (=1′b0), to the left by 1 bit is set as the variable “code”.
328 328 328 Furthermore, for example, in a second while loop, the code assignment unitselects, as a target to which a code bit string is to be assigned, a symbol that has the second shortest code length, or a symbol has the shortest code length and is the second in the specific symbol order when there is another symbol having the same code length. The code assignment unitassigns the variable “code” (=2′b10) as a code bit string of the selected symbol. Then, the code assignment unitperforms a shift operation of shifting a value that is obtained by adding one to the variable “code”, to the left by the number of bits obtained by subtracting the code length of the current symbol from the code length of the next symbol, and sets the value obtained by the shift operation as the variable “code”. For example, in a case where the number of bits obtained by subtracting the code length of the current symbol from the code length of the next symbol is one, a value 3′b110 obtained by shifting 2′b11 that is a value obtained by adding one to the variable “code” (=2′b10), to the left by 1 bit is set as the variable “code”.
328 328 By performing such loop processing, the code assignment unitcan assign a code bit string to each of the symbols by using the relationship in order between the symbols and using the code lengths corresponding to the respective symbols. In other words, in a case where the relationship in order between the symbols and the code lengths of the respective symbols have been determined, the code assignment unitcan uniquely determine a code bit string to be assigned to each of the symbols.
33 This indicates that, in a case where compressed data output by the variable length coding unitincludes, as a header, the code lengths that have been arranged in predetermined order and encoded, a code table (decoding table) to be used for decoding in the decompressor can be restored from the code lengths. By encoding the code lengths arranged in order of the symbols, a code amount overhead of the compressed data can be greatly reduced as compared with a case of encoding the code bit strings themselves.
328 33 The code assignment unitsends the code table indicative of the code bit string assigned to each symbol to the variable length coding unit.
33 31 32 As described above, the variable length coding unitconverts the symbols sent from the buffer unitinto variable length codes by using the code table sent by the code table generation unit, and outputs the variable length codes as compressed data.
15 326 326 326 326 With the above configuration, the variable length coding devicecan reduce the processing time and the processing resources when assigning a variable length code that has a code length shorter than or equal to the maximum code length and is a perfect code, to a symbol. Specifically, in the maximum code length restriction process, the maximum code length restriction unitrestricts a code length of each symbol to a code length shorter than or equal to the maximum code length while maintaining Huffman codes to be assigned based on the code length of each symbol being perfect codes. Therefore, the maximum code length restriction unitcan guarantee that the Huffman codes to be assigned based on the code length of each symbol are perfect codes when the maximum code length restriction process has been completed. Therefore, the maximum code length restriction unitcan reduce the processing time and the processing resources as compared to, for example, the maximum code length restriction unitA of the comparative example, which additionally executes a process of merging a binary tree having a redundant leaf node.
326 When the code length of the representative symbol is longer than the representative symbol maximum code length, the maximum code length restriction unitof the first embodiment swaps (exchanges) the code lengths between the representative symbol and a non-representative symbol that satisfies a condition that the code length of the non-representative symbol is shorter than or equal to the representative symbol maximum code length.
326 However, when the code length of the representative symbol is longer than the representative symbol maximum code length, there is no non-representative symbol that satisfies the condition that its code length is shorter than or equal to the representative symbol maximum code length, in some cases. A maximum code length restriction unitaccording to a second embodiment is further configured to change the code lengths of the representative and non-representative symbols in order to restrict (limit) the code length of the representative symbol to the representative symbol maximum code length in such cases.
15 15 15 15 A configuration of the variable length coding deviceaccording to the second embodiment is similar to that of the variable length coding deviceof the first embodiment. The variable length coding deviceof the second embodiment is different from the variable length coding deviceof the first embodiment in terms of a configuration that changes the code lengths of the representative and non-representative symbols being added in order to restrict the code length of the representative symbol to the representative symbol maximum code length in the above-described cases. Hereinafter, the difference from the first embodiment will be mainly described.
27 FIG. 28 FIG. A case where when the code length of the representative symbol is longer than the representative symbol maximum code length, there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length will be specifically described with reference toand.
27 FIG. 27 FIG. 323 15 shows an example of non-representative symbols and a representative symbol determined based on frequencies of occurrence by the symbol merge unitof the variable length coding device. In the example shown in, symbols are arranged in descending order of the frequencies of occurrence. Specifically, a frequency of occurrence of each of symbols “g” and “f” is eight. A frequency of occurrence of each of symbols “e” and “d” is seven. A frequency of occurrence of each of symbols “c” and “b” is six. A frequency of occurrence of a symbol “a” is five. A frequency of occurrence of each of symbols “h”, “i”, “j”, and “k” is one.
The bottom four symbols “h”, “i”, “j”, and “k are represented by one representative symbol R among the symbols in descending order of the frequencies of occurrence. In other words, the four symbols “h”, “i”, “j”, and “k” are merge symbols. The sum of the frequencies of occurrence of the four merge symbols “h”, “i”, “j”, and “k”, i.e., four is set as a frequency of occurrence of the representative symbol “R”. The remaining seven symbols “a”, “b”, “c”, “d”, “e”, “f”, and “g” are non-representative symbols (hereinafter referred to as non-representative symbols a” to “g”).
28 FIG. 27 FIG. 8 324 15 8 shows an example of a Huffman treeA generated by the Huffman tree generation unitof the variable length coding device. The Huffman treeA corresponds to the non-representative symbols and the representative symbol shown in. Here, it is assumed that the maximum code length is 4 bits.
8 8 11 3 FIG. The Huffman treeA is generated based on the frequency of occurrence of each of the seven non-representative symbols “a” to “g” and the frequency of occurrence of the representative symbol “R”. Specifically, the Huffman treeA is generated by, for example, the Huffman tree generation process described above with reference toin which step Sis replaced with a step of adding, as leaf nodes, the seven non-representative symbols “a” to “g” and the representative symbol “R”.
8 In the generated Huffman treeA, a code length of each of the non-representative symbols “a” to “g” is 3 bits. In addition, a code length of the representative symbol “R” is 3 bits.
8 8 Note that the edges and nodes represented by dotted lines indicate a subtreeB corresponding to the merge symbols “h”, “i”, “j”, and “k” (hereinafter referred to as merge symbols “h” to “k”). The subtreeB is a subtree having the representative symbol “R” as its root node and having each of the merge symbols “h” to “k” as a leaf node. A code length of each of the merge symbols “h” to “k” is 5 bits. Therefore, for these merge symbols “h” to “k”, the maximum value of the code lengths which are to be added to the code length of the representative symbol “R” (i.e., the maximum value of additional code lengths) is two.
8 373 8 In the Huffman treeA, a length obtained by adding the maximum value of the additional code lengths to the code length of the representative symbol “R” is longer than the maximum code length (i.e., 3+2 >4). Accordingly, the representative symbol swapping unitattempts to select a non-representative symbol having a code length shorter than or equal to the representative symbol maximum code length that is obtained by subtracting the maximum value of the additional code lengths from the maximum code length (=4−2=2). In the Huffman treeA, however, there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length.
8 Therefore, the Huffman treeA corresponds to the case where when the code length of the representative symbol is longer than the representative symbol maximum code length, there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length.
326 15 354 325 326 The maximum code length restriction unitof the variable length coding deviceexecutes a process of restricting (limiting) code lengths corresponding to respective symbols to the maximum code length by using the maximum value of additional code lengths sent from the merge symbol additional code length determination unit, and H <non-representative symbol, code length> pairs and one <representative symbol, code length> pair sent from the code length determination unit(i.e., maximum code length restriction process). That is, the maximum code length restriction unitexecutes the maximum code length restriction process by using the H <non-representative symbol, code length> pairs and one <representative symbol, code length, maximum value of additional code lengths> pair.
29 FIG. 326 326 371 372 373 381 382 383 384 374 375 376 377 is a block diagram showing an example of a configuration of the maximum code length restriction unit. The maximum code length restriction unitincludes, for example, a code length sorting unit, a code length clipping unit, a representative symbol swapping unit, a second representative symbol swapping unit, a sibling symbol code length change unit, a representative symbol code length change unit, a representative symbol restriction process termination determination unit, a violation symbol count calculation unit, a termination determination unit, a code length change unit, and a violation symbol count decrement unit.
371 372 373 374 375 376 377 373 18 FIG. The operations of the code length sorting unit, the code length clipping unit, the representative symbol swapping unit, the violation symbol count calculation unit, the termination determination unit, the code length change unit, and the violation symbol count decrement unithave been described above with reference to. More specifically, the operation in a case where the representative symbol swapping unitswaps the code lengths between the representative symbol and a non-representative symbol which is shorter than the representative symbol maximum code length, has been described as the maximum code length restriction process in the first embodiment.
373 When the code length of the representative symbol is longer than the representative symbol maximum code length, the representative symbol swapping unitdoes not swap the code lengths between the representative symbol and a non-representative symbol unless there is a non-representative symbol having a code length shorter than or equal to the representative symbol maximum code length. Hereinafter, an operation in a case where when the code length of the representative symbol is longer than the representative symbol maximum code length, there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length will be described.
381 381 When the code length of the representative symbol is longer than the representative symbol maximum code length, the second representative symbol swapping unitswaps the code length of the representative symbol for the shortest code length among the code lengths of the non-representative symbols shorter than the code length of the representative symbol (i.e., swaps the code lengths between the representative symbol and the non-representative symbol). Note that when the code length of the representative symbol is longer than the representative symbol maximum code length but there is no non-representative symbol whose code length is shorter than the code length of the representative symbol, the second representative symbol swapping unitdoes not swap the code lengths between the representative symbol and a non-representative symbol.
382 383 384 381 The sibling symbol code length change unit, the representative symbol code length change unit, and the representative symbol restriction process termination determination unitexecute a process of restricting the code length of the representative symbol to the representative symbol maximum code length (representative symbol restriction process) after the process by the second representative symbol swapping unitis executed.
382 382 326 D D D 0 First, the sibling symbol code length change unitselects 2non-representative symbols s that have, as their parent, the same intermediate node as the representative symbol or that have, as their transitive parent, the same intermediate node as the representative symbol. The non-representative symbols that have, as their parent, the same intermediate node as the representative symbol or that have, as their transitive parent, the same intermediate node as the representative symbol are also referred to as sibling symbols. The transitive parent is a node which can be reached while edges are traced from a symbol (leaf node) to the root node. D is a difference obtained by subtracting the code length of the representative symbol from the code length of the selected sibling symbol. In other words, the code lengths of the selected sibling symbols are the code length of the representative symbol+D. Specifically, when 2sibling symbols s of the representative symbol are searched for, it is sufficient for the sibling symbol code length change unitto select 2non-representative symbols each having a code length obtained by adding D to the code length of the representative symbol, from the H <non-representative symbol, code length> pairs. Note that when a non-representative symbol having, as its parent, the same intermediate node as the representative symbol is selected, the maximum code length restriction unitselects one (=2) sibling symbol s since there is no difference between the code length of the representative symbol and the code length of the selected non-representative symbol.
382 382 382 382 D D D D D Next, the sibling symbol code length change unitselects 2non-representative symbols s′ that have code lengths shorter than the maximum code length, in the order of longer code lengths, from the non-representative symbols. Then, the sibling symbol code length change unitchanges the code length of each of the 2sibling symbols s and the code length of each of the 2non-representative symbols s′. For example, a case where a code length of an i-th sibling symbol s among the 2sibling symbols s and a code length of an i-th non-representative symbol s′ among the 2non-representative symbols s′ will be explained. Note that i is any value between one to D inclusive. In this case, the sibling symbol code length change unitchanges the code length of the i-th sibling symbol s to a length obtained by adding one to the code length of the i-th non-representative symbol s′. Then, the sibling symbol code length change unitincrements the code length of the i-th non-representative symbol s′ by one.
383 383 Next, the representative symbol code length change unitdecrements the code length of the representative symbol by one. In other words, the representative symbol code length change unitchanges the code length of the <representative symbol, code length, maximum value of additional code lengths> pair by decrementing the code length by one.
384 384 The representative symbol restriction process termination determination unitdetermines whether or not to terminate the representative symbol restriction process, based on the decremented code length of the representative symbol. Specifically, the representative symbol restriction process termination determination unitdetermines whether or not the decremented code length of the representative symbol becomes shorter than or equal to the representative symbol maximum code length.
382 383 384 When the decremented code length of the representative symbol is still longer than the representative symbol maximum code length, the representative symbol restriction process by the sibling symbol code length change unit, the representative symbol code length change unit, and the representative symbol restriction process termination determination unitdescribed above is executed again. In other words, the representative symbol restriction process is repeated until the code length of the representative symbol becomes shorter than or equal to the representative symbol maximum code length.
384 When the decremented code length of the representative symbol becomes shorter than or equal to the representative symbol maximum code length, the representative symbol restriction process termination determination unitends the representative symbol restriction process.
374 375 376 377 The subsequent process of the violation symbol count calculation unit, the termination determination unit, the code length change unit, and the violation symbol count decrement unithas been described above in the maximum code length restriction process of the first embodiment.
326 326 With the above configuration, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length, even in a case where there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length. Furthermore, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length even when there is no non-representative symbol corresponding to a code length shorter than the code length of the representative symbol.
326 30 FIG. 34 FIG. A specific example in which code lengths are restricted to the maximum code length in the maximum code length restriction unitwill be described with reference toto. Here, it is assumed that the maximum code length is 4 bits.
30 FIG. 326 shows an example of a Huffman tree generated in the maximum code length restriction unit.
80 832 833 834 835 836 837 838 831 A Huffman treeis generated based on a frequency of occurrence of each of seven non-representative symbols “a”, “b”, “c”, “d”, “e”, “f”, and “g” (hereinafter referred to as non-representative symbols “a” to “g”) and a frequency of occurrence of one representative symbol “R”. The non-representative symbols “a” to “g” are assigned to leaf nodes,,,,,, and, respectively. The representative symbol “R” is assigned to a leaf node.
80 In the generated Huffman tree, code lengths of the non-representative symbols “a” to “g” are 3 bits. In addition, a code length of the representative symbol “R” is 3 bits.
85 85 831 861 862 863 864 865 866 867 868 Note that the edges and nodes represented by dotted lines indicate a subtreecorresponding to merge symbols “h”, “i”, “j”, “k”, “l”, “m”, “n”, and “o” (hereinafter referred to as merge symbols “h” to “o”). The subtreeis a subtree having the representative symbol “R” (leaf node) as its root node and having each of the merge symbols “h” to “o” as a leaf node. The merge symbols “h” to “o” are assigned to leaf nodes,,,,,,, and, respectively. The code lengths of the merge symbols “h” to “o” are 6 bits. Therefore, for these merge symbols “h” to “o”, the maximum value of the code lengths which are to be added to the code length of the representative symbol “R” (i.e., the maximum value of the additional code lengths) is three.
80 373 80 Furthermore, in the Huffman tree, a length obtained by adding the maximum value of the additional code lengths to the code length of the representative symbol “R” is longer than the maximum code length (i.e., 3+3 >4). Accordingly, the representative symbol swapping unitattempts to select a non-representative symbol whose code length is shorter than or equal to a representative symbol maximum code length obtained by subtracting the maximum value of the additional code lengths from the maximum code length (=4−3=1). In the Huffman tree, however, there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length.
381 80 In this case, the second representative symbol swapping unitattempts to select a non-representative symbol whose code length is shorter than the code length of the representative symbol “R”. In the Huffman tree, however, there is also no non-representative symbol satisfying the condition that its code length is shorter than the code length of the representative symbol (for example, there is no non-representational symbol whose code length is 2 bits).
382 821 382 30 FIG. 30 FIG. Accordingly, the sibling symbol code length change unitselects a non-representative symbol that has, as its parent, the same intermediate nodeas the representative symbol “R”. In the example shown in, the non-representative symbol “a” is selected. Then, the sibling symbol code length modification unitselects a non-representative symbol having the longest code length among the non-representative symbols whose code lengths are shorter than the maximum code length. In the example shown in, the non-representative symbol “g” is selected.
382 382 The sibling symbol code length change unitchanges the code length of the non-representative symbol “a” to a length obtained by adding one to the code length of the non-representative symbol “g” (=3+1=4 bits). Then, the sibling symbol code length change unitincrements the code length of the non-representative symbol “g” by one, thereby changing the code length to 4 bits.
31 FIG. 30 FIG. 80 80 382 80 843 844 838 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the sibling symbol code length change unithas changed the code lengths of the non-representative symbols “a” and “g” as described above. That is, in the modified Huffman tree, the code lengths of the non-representative symbols “a” and “g” are 4 bits. The non-representative symbols “a” and “g” are assigned to leaf nodesand, respectively, that have, as their parent, the same intermediate node.
832 843 821 831 821 831 383 Since the code length of the non-representative symbol “a” has been changed (i.e., since the leaf node to which the non-representative symbol “a” is assigned has been changed from the leaf nodeto the leaf node), leaf nodes having the intermediate nodeas their parent is only the leaf nodeof the representative symbol “R”. In this case, the intermediate nodeand the leaf nodecan be merged. Thus, the representative symbol code length change unitdecrements the code length of the representative symbol “R” by one, thereby changing the code length to 2 bits.
32 FIG. 31 FIG. 80 80 383 80 821 821 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the representative symbol code length change unithas changed the code length of the representative symbol “R” as described above. That is, in the modified Huffman tree, the code length of the representative symbol “R” is set to 2 bits. The representative symbol “R” is assigned to the leaf node, which has been changed from the parent intermediate node.
384 384 Next, the representative symbol restriction process termination determination unitdetermines whether or not the code length of the representative symbol “R” is shorter than or equal to the representative symbol maximum code length. Since the code length of the representative symbol “R” (=2) is longer than the representative symbol maximum code length (=1), the representative symbol restriction process termination determination unitdecides to change code lengths of sibling symbols and the code length of the representative symbol again.
382 80 382 833 834 811 32 FIG. 1 The sibling symbol code length change unitattempts to select a non-representative symbol that has, as its parent, the same intermediate node as the representative symbol “R”. In the Huffman tree, however, there is no non-representative symbol that has, as its parent, the same intermediate node as the representative symbol “R”. In this case, the sibling symbol code length change unitselects non-representative symbols that have, as their transitive parent, the same intermediate node as the representative symbol “R”. In, the two non-representative symbols “b” and “c” (leaf nodesand) that have, as their transitive parent, the intermediate nodewhich is the parent of the representative symbol “R”, are selected. In other words, the two (=2) non-representative symbols “b” and “c” each having a code length obtained by adding one to the code length of the representative symbol “R” are selected.
382 32 FIG. Next, the sibling symbol code length change unitselects two non-representative symbols having code lengths shorter than the maximum code length, in the order of longer code lengths, from the non-representative symbols. In, the non-representative symbols “f” and “e” are selected.
382 382 The sibling symbol code length change unitchanges the code length of the selected non-representative symbol “b” to a length obtained by adding one to the code length of the non-representative symbol “f” (=3+1=4 bits). Then, the sibling symbol code length change unitincrements the code length of the non-representative symbol “f” by one, thereby changing the code length to 4 bits.
382 382 In addition, the sibling symbol code length change unitchanges the code length of the selected non-representative symbol “c” to a length obtained by adding one to the code length of the non-representative symbol “e” (=3+1=4 bits). Then, the sibling symbol code length change unitincrements the code length of the non-representative symbol “e” by one, thereby changing the code length to 4 bits.
33 FIG. 32 FIG. 80 80 382 80 849 84 837 847 848 836 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the sibling symbol code length change unithas changed the code lengths of the non-representative symbols “b”, “f”, “c”, and “e” as described above. That is, in the modified Huffman tree, the code lengths of the non-representative symbols “b”, “f”, “c”, and “e” are set to 4 bits. The non-representative symbols “f” and “b” are assigned to respective leaf nodesandA that have the intermediate nodeas their parent. The non-representative symbols “e” and “c” are assigned to respective leaf nodesandthat have the intermediate nodeas their parent.
833 84 834 848 811 821 822 821 811 383 Since the code lengths of the non-representative symbols “b” and “c” have been changed (i.e., since the leaf node to which the non-representative symbol “b” is assigned has been changed from the leaf nodeto the leaf nodeA and the leaf node to which the non-representative symbol “c” is assigned has been changed from the leaf nodeto the leaf node), leaf nodes that have the intermediate nodeas their parent is only the leaf nodeof the representative symbol “R”. In this case, a binary tree that includes the redundant intermediate node, the leaf node, and the intermediate nodecan be merged. Accordingly, the representative symbol code length change unitdecrements the code length of the representative symbol “R” by one, thereby changing the code length to 1 bit.
34 FIG. 33 FIG. 80 80 383 80 811 811 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the representative symbol code length change unithas changed the code length of the representative symbol “R” as described above. That is, in the modified Huffman tree, the code length of the representative symbol “R” is set to 1 bit. The representative symbol “R” is assigned to the leaf node, which has been changed from the parent intermediate node.
384 384 Next, the representative symbol restriction process termination determination unitdetermines whether or not the code length of the representative symbol “R” is shorter than or equal to the representative symbol maximum code length. Since the code length of the representative symbol “R” (=1) is shorter than or equal to the representative symbol maximum code length (=1), the representative symbol restriction process termination determination unitdecides to terminate the process of changing the code lengths of the representative symbol and sibling symbols.
80 80 80 80 34 FIG. The Huffman treeshown inis a valid Huffman tree and every intermediate node of the Huffman treehas two child nodes. Therefore, Huffman codes determined based on the Huffman treeare Huffman codes that satisfy Kraft's inequality and have no redundant code assignment. That is, the Huffman codes determined based on the Huffman treeare perfect codes.
35 FIG. 36 FIG. An example of the procedure of the maximum code length restriction process will be described with reference toand.
35 FIG. 326 326 325 is a flowchart showing an example of the procedure of the maximum code length restriction process executed in the maximum code length restriction unit. The maximum code length restriction unitexecutes the maximum code length restriction process on H <non-representative symbol, code length> pairs and one <representative symbol, code length> pair that are sent from the code length determination unit.
326 501 326 502 First, the maximum code length restriction unitchanges the code lengths of the non-representative symbols that are longer than the maximum code length, to the maximum code length (step S). Next, the maximum code length restriction unitdetermines whether or not the code length of the representative symbol is longer than the representative symbol maximum code length (step S).
502 326 507 When the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length (No in step S), the process by the maximum code length restriction unitproceeds to step S.
502 326 503 When the code length of the representative symbol is longer than the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitdetermines whether or not there is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (step S).
503 326 505 326 506 507 When there is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitselects a non-representative symbol having the longest code length among the non-representative symbols each having a code length shorter than or equal to the representative symbol maximum code length (step S). The maximum code length restriction unitswaps the code lengths between the selected non-representative symbol and the representative symbol (step S) and proceeds to step S.
503 326 504 507 36 FIG. When there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (No in step S), the maximum code length restriction unitexecutes a swap and code length change process (step S) and proceeds to step S. The swap and code length change process is a process for restricting the code length of the representative symbol to the representative symbol maximum code length by: swapping the code lengths between the representative symbol and a non-representative symbol; changing a code length of a sibling symbol (or code lengths of sibling symbols) of the representative symbol; and changing the code length of the representative symbol. A specific procedure of the swap and code length change process will be described later with reference to a flowchart in.
507 510 305 308 326 24 FIG. The subsequent procedure from step Sto step Sis the same as the procedure from step Sto step Sof the maximum code length restriction process described above with reference to. In other words, the maximum code length restriction unitexecutes a process for changing a code length of a violation symbol to a code length shorter than or equal to the maximum code length.
326 With the above maximum code length restriction process, the maximum code length restriction unitcan restrict the code lengths of the non-representative symbols to the maximum code length and restrict the code length of the representative symbol to the representative symbol maximum code length. By restricting the code length of the representative symbol to the representative symbol maximum code length, code lengths of symbols represented by the representative symbol (i.e., code lengths of merge symbols) can be restricted to the maximum code length.
36 FIG. 35 FIG. 326 504 326 is a flowchart showing an example of the procedure of the swap and code length change process executed in the maximum code length restriction unit. The swap and code length change process corresponds to step Sof the maximum code length restriction process described above with reference to. The maximum code length restriction unitexecutes the swap and code length change process when there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length.
326 601 First, the maximum code length restriction unitdetermines whether or not there is a non-representative symbol having a code length shorter than the code length of the representative symbol (step S).
601 326 604 When there is no non-representative symbol having a code length shorter than the code length of the representative symbol (No in step S), the process by the maximum code length restriction unitproceeds to step S.
601 326 602 326 603 604 When there is a non-representative symbol having a code length shorter than the code length of the representative symbol (Yes in step S), the maximum code length restriction unitselects a non-representative symbol having the shortest code length, among the non-representative symbols each having a code length shorter than the code length of the representative symbol (step S). The maximum code length restriction unitswaps the code lengths between the selected non-representative symbol and the representative symbol (step S) and proceeds to step S.
326 604 D D Next, the maximum code length restriction unitselects 2non-representative symbols s (i.e., 2sibling symbols s) that have, as their parent, the same intermediate node as the representative symbol or that have, as their transitive parent, the same intermediate node as the representative symbol (step S).
326 605 326 606 D D D The maximum code length restriction unitselects 2non-representative symbols s′ each having a code length shorter than the maximum code length, in the order of longer code lengths, from the non-representative symbols (step S). Then, the maximum code length restriction unitsets a variable i to one (step S). The variable i is a variable for specifying the i-th sibling symbol s of the 2sibling symbols s from the head and the i-th non-representative symbol s′ of the 2non-representative symbols s′ from the head.
326 607 326 608 326 609 326 610 D The maximum code length restriction unitchanges the code length of the i-th sibling symbol s to a code length obtained by adding one to the code length of the i-th non-representative symbol s′ (step S). The maximum code length restriction unitincrements the code length of the i-th non-representative symbol s′ by one (step S). The maximum code length restriction unitincrements the variable i by one (step S). Then, the maximum code length restriction unitdetermines whether or not the variable i is larger than 2(step S).
D D D 610 326 607 326 607 609 When the variable i is smaller than or equal to 2(No in step S), the maximum code length restriction unitreturns to step Sand further executes the steps to change the code length of the i-th sibling symbol s and the code length of the i-th non-representative symbol s′. In other words, the maximum code length restriction unitrepeats the procedure from step Sto step Suntil all the code lengths of the selected 2sibling symbols s and 2non-representative symbols s′ are changed.
D 610 326 611 326 612 When the variable i is larger than 2(Yes in step S), the maximum code length restriction unitdecrements the code length of the representative symbol by one (step S). Then, the maximum code length restriction unitdetermines whether or not the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length (step S).
612 326 604 326 When the code length of the representative symbol is longer than the representative symbol maximum code length (No in step S), the process by the maximum code length restriction unitreturns to step S. That is, the maximum code length restriction unitfurther executes the steps of changing code lengths of one or more sibling symbols of the representative symbol in order to shorten the code length of the representative symbol.
612 326 When the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitends the swap and code length change process.
326 326 With the above swap and code length change process, even when there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length. Furthermore, even when there is no non-representative symbol whose code length is shorter than the code length of the representative symbol, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length.
326 326 Similarly to the maximum code length restriction unitof the second embodiment, a maximum code length restriction unitof a third embodiment is configured to change a code length of a representative symbol and code lengths of non-representative symbols in order to restrict the code length of the representative symbol to a representative symbol maximum code length in a case where there is no non-representative symbol satisfying a condition that its code length is shorter than or equal to the representative symbol maximum code length when the code length of the representative symbol is longer than the representative symbol maximum code length.
15 15 15 15 A configuration of a variable length coding deviceaccording to the third embodiment is similar to that of the variable length coding devicesof the first and second embodiments. The variable length coding deviceof the third embodiment is different from the variable length coding devicesof the first and second embodiments in terms of the procedure of changing the code lengths of the representative and non-representative symbols in order to restrict the code length of the representative symbol to the representative symbol maximum code length in the above-described case. Hereinafter, the difference from the second embodiment will be mainly described.
326 15 354 325 326 The maximum code length restriction unitof the variable length coding deviceexecutes a process of restricting code lengths corresponding to respective symbols to the maximum code length (maximum code length restriction process) by using the maximum value of additional code lengths sent from the merge symbol additional code length determination unit, H <non-representative symbol, code length> pairs and one <representative symbol, code length> pair that are sent from the code length determination unit. In other words, the maximum code length restriction unitexecutes the maximum code length restriction process by using the H <non-representative symbol, code length> pairs and one <representative symbol, code length, maximum value of additional code lengths> pair.
37 FIG. 326 326 371 372 373 391 392 393 374 375 376 377 is a block diagram showing an example of a configuration of the maximum code length restriction unit. The maximum code length restriction unitincludes, for example, a code length sorting unit, a code length clipping unit, a representative symbol swapping unit, an intermediate node selection unit, a transitive leaf code length change unit, a representative symbol code length change unit, a violation symbol count calculation unit, a termination determination unit, a code length change unit, and a violation symbol count decrement unit.
371 372 373 374 375 376 377 373 18 FIG. The operations of the code length sorting unit, the code length clipping unit, the representative symbol swapping unit, the violation symbol count calculation unit, the termination determination unit, the code length change unit, and the violation symbol count decrement unithave been described above with reference to. More specifically, the operation in which the representative symbol swapping unitswaps the code lengths between the representative symbol and the non-representative symbol that is shorter than the representative symbol maximum code length has been described in the maximum code length restriction process of the first embodiment.
373 When the code length of the representative symbol is longer than the representative symbol maximum code length, the representative symbol swapping unitdoes not swap the code lengths between the representative symbol and a non-representative symbol unless there is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length. Hereinafter, an operation in a case where there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length when the code length of the representative symbol is longer than the representative symbol maximum code length will be described.
373 391 392 393 After the process of the representative symbol swapping unitis executed, the intermediate node selection unit, the transitive leaf code length change unit, and the representative symbol code length change unitexecute a process of restricting the code length of the representative symbol to the representative symbol maximum code length (representative symbol restriction process).
391 324 First, the intermediate node selection unitselects one intermediate node whose depth from the root node is equal to the representative symbol maximum code length, in a Huffman tree generated by the Huffman tree generation unit.
392 The transitive leaf code length change unitidentifies all leaf nodes that can be reached by tracing from the selected intermediate node in a direction of child nodes. The leaf nodes that can be reached by tracing from the intermediate node in a direction of child nodes are referred to as transitive leaf nodes. The direction of child nodes is a deeper direction.
392 392 The transitive leaf code length modification unitchanges a code length of one non-representative symbol among the non-representative symbols assigned to the identified transitive leaf nodes, to the code length of the representative symbol. The transitive leaf code length change unitchanges the code lengths of the remaining non-representative symbols among the non-representative symbols assigned to the transitive leaf nodes, to the maximum code length.
393 Then, the representative symbol code length change unitchanges the code length of the representative symbol to the representative symbol maximum code length.
374 375 376 377 The subsequent operations of the violation symbol count calculation unit, the termination determination unit, the code length change unit, and the violation symbol count decrement unithave been described above as the maximum code length restriction process of the first embodiment.
326 With the above configuration, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length, even in a case where there is no non-representative symbol satisfying the condition that its code length is shorter than or equal to the representative symbol maximum code length.
326 38 FIG. 40 FIG. A specific example in which code lengths are restricted to the maximum code length in the maximum code length restriction unitwill be described with reference toto. Here, it is assumed that the maximum code length is 4 bits.
38 FIG. 326 shows an example of a Huffman tree generated in the maximum code length restriction unit.
90 932 933 934 935 936 937 938 931 90 The Huffman treeis generated based on a frequency of occurrence of each of seven non-representative symbols “a”, “b”, “c”, “d”, “e”, “f”, and “g” (hereinafter referred to as non-representative symbols “a” to “g”) and a frequency of occurrence of one representative symbol “R”. The non-representative symbols “a” to “g” are assigned to leaf nodes,,,,,, and, respectively. The representative symbol “R” is assigned to a leaf node. In the generated Huffman tree, code lengths of the non-representative symbols “a” to “g” are 3 bits. In addition, a code length of the representative symbol “R” is 3 bits.
95 95 931 951 952 953 954 Note that the edges and nodes represented by dotted lines indicate a subtreecorresponding to merge symbols “h”, “i”, “j”, and “k” (hereinafter referred to as merge symbols “h” to “k”). The subtreeis a subtree that has the representative symbol “R” (leaf node) as its root node and has each of the merge symbols “h” to “k” as a leaf node. The merge symbols “h” to “k” are assigned to leaf nodes,,, and, respectively. The code lengths of the merge symbols “h” to “k” are 5 bits. Therefore, for these merge symbols “h” to “k”, the maximum value of code lengths which are to be added to the code length of the representative symbol “R” (i.e., the maximum value of additional code lengths) is two (=5−3).
90 373 90 Furthermore, in the Huffman tree, a length obtained by adding the maximum value of the additional code lengths to the code length of the representative symbol “R” is longer than the maximum code length (i.e., 3+2 >4). Accordingly, the representative symbol swapping unitattempts to select a non-representative symbol whose code length is shorter than or equal to a representative symbol maximum code length obtained by subtracting the maximum value of the additional code lengths from the maximum code length (=4−2=2). In the Huffman tree, however, there is no non-representative symbol satisfying a condition that its code length is shorter than or equal to the representative symbol maximum code length.
391 901 924 901 391 931 911 921 922 391 38 FIG. 38 FIG. Therefore, the intermediate node selection unitselects an intermediate node whose depth from a root nodeis equal to the representative symbol maximum code length. In, an intermediate nodewhose depth from the root nodeis two, is selected. Note that the intermediate node selection unitexcludes, from targets of the selection, intermediate nodes each of which is a parent or a transitive parent of the leaf nodeto which the representative symbol “R” is assigned. In, intermediate nodes,, andare excluded from the targets from which the intermediate node selection unitselects an intermediate node.
392 924 937 938 38 FIG. The transitive leaf code length change unitidentifies all leaf nodes (i.e., transitive leaf nodes) that can be reached by tracing from the selected intermediate nodein a direction of child nodes. In, a leaf nodeand a leaf nodeare identified.
392 937 938 38 FIG. The transitive leaf code length change unitselects one non-representative symbol from the non-representative symbols “f” and “g” that are assigned to the identified leaf nodeand leaf node, respectively. In, the non-representative symbol “f” is selected.
392 392 The transitive leaf code length change unitchanges the code length of the selected non-representative symbol “f” to the code length of the representative symbol “R” (=3 bits). In addition, the transitive leaf code length change unitchanges the code lengths of all the remaining non-representative symbols assigned to the transitive leaf nodes (in this case, the code length of the non-representative symbol “g”), to the maximum code length.
393 924 Then, the representative symbol code length change unitchanges the code length of the representative symbol “R” to the representative symbol maximum code length (i.e., the depth of the selected intermediate node).
39 FIG. 38 FIG. 90 90 392 393 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the transitive leaf code length change unithas been changed the code lengths of the non-representative symbols “f” and “g” and the representative symbol code length change unithas been changed the code length of the representative symbol “R” as described above.
90 In the modified Huffman tree, the code length of the non-representative symbol “f” is set to the original code length of the representative symbol “R” (=3 bits). The code length of the representative symbol “R” is set to the representative symbol maximum code length (=2 bits). The code length of the non-representative symbol “g” is set to the maximum code length (=4 bits).
943 944 945 946 95 947 901 947 Since the code length of the representative symbol “R” is set to the representative symbol maximum code length, code lengths of the merge symbols “h” to “k” that are assigned to leaf nodes,,, andof the subtree, respectively, become shorter than or equal to the maximum code length. However, since a leaf nodeis a node that cannot be traced from the root node, the non-representative symbol “g” assigned to the leaf nodeis a violation symbol.
374 In this case, the violation symbol count calculation unitcalculates the number I of violation symbols by using the above-described equation (7) (or equation (8)).
376 Since the calculated number I of violation symbols is larger than 0, the code length change unitchanges the code lengths of the non-representative symbols until the number I of violation symbols becomes 0.
376 376 376 Specifically, the code length change unitselects the violation symbol “g”. Then, the code length change unitselects the non-representative symbol “f” having a code length shorter than the maximum code length. Note that if there are multiple non-representative symbols each having a code length shorter than the maximum code length, the code length change unitselects a non-representative symbol from the multiple non-representative symbols in the order of longer code lengths.
376 376 The code length change unitchanges the code length of the selected violation symbol “g” to a length obtained by adding one to the code length of the selected non-representative symbol “f” (=3+1=4 bits). Then, the code length change unitincrements the code length of the non-representative symbol “f” by one, thereby changing the code length to 4 bits.
377 375 Next, the violation symbol count decrement unitdecrements the number I of violation symbols by one, thereby updating the number I of violation symbols with 0. Since the updated number I of violation symbols is 0, the termination determination unitdecides to terminate the change of the code lengths of the non-representative symbols.
40 FIG. 39 FIG. 90 90 376 90 90 shows a modified example of the Huffman treeshown in. This modified example shows the Huffman treein which the code length change unithas changed the code length of the non-representative symbol “f” and the code length of the violation symbol “g”. That is, in the modified Huffman tree, the code length of the non-representative symbol “f” and the code length of the violation symbol “g” are set to 4 bits. In addition, this Huffman treehas no violation symbols.
90 90 90 90 The Huffman treeis a valid Huffman tree and every intermediate node of the Huffman treehas two child nodes. Therefore, Huffman codes determined based on the Huffman treeare Huffman codes that satisfy Kraft's inequality and have no redundant code assignment (i.e., K=1). In other words, the Huffman codes determined based on the Huffman treeare perfect codes.
41 FIG. 42 FIG. An example of the procedure of the maximum code length restriction process will be described with reference toand.
41 FIG. 326 326 325 is a flowchart showing an example of the procedure of the maximum code length restriction process executed in the maximum code length restriction unit. The maximum code length restriction unitexecutes the maximum code length restriction process on H <non-representative symbol, code length> pairs and one <representative symbol, code length> pair, which are sent from the code length determination unit.
326 701 326 702 First, the maximum code length restriction unitchanges the code lengths of the non-representative symbols that are longer than the maximum code length, to the maximum code length (step S). Next, the maximum code length restriction unitdetermines whether or not the code length of the representative symbol is longer than the representative symbol maximum code length (step S).
702 326 707 When the code length of the representative symbol is shorter than or equal to the representative symbol maximum code length (No in step S), the process by the maximum code length restriction unitproceeds to step S.
702 326 703 When the code length of the representative symbol is longer than the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitdetermines whether or not there is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (step S).
703 326 705 326 706 707 When there is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (Yes in step S), the maximum code length restriction unitselects a non-representative symbol having the longest code length, among the non-representative symbols each having a code length shorter than or equal to the representative symbol maximum code length (step S). The maximum code length restriction unitswaps the code lengths between the selected non-representative symbol and the representative symbol (step S) and proceeds to step S.
703 326 704 707 42 FIG. When there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (No in step S), the maximum code length restriction unitexecutes a second code length change process (step S) and proceeds to step S. The second code length change process is a process for changing the code length of the representative symbol to a code length shorter than or equal to the representative symbol maximum code length by: identifying an intermediate node having a depth from the root node that is equivalent to the representative symbol maximum code length, changing code lengths of non-representative symbol to which transitive leaf nodes of this intermediate node are assigned, respectively, and changing the code length of the representative symbol. A specific procedure of the second code length change process will be described later with reference to a flowchart in.
707 710 305 308 326 24 FIG. The subsequent procedure from step Sto step Sis similar to the procedure from step Sto step Sof the maximum code length restriction process described above with reference to. That is, the maximum code length restriction unitexecutes a process of changing a code length of a violation symbol to a code length shorter than or equal to the maximum code length.
326 With the above maximum code length restriction process, the maximum code length restriction unitcan restrict the code lengths of the non-representative symbols to the maximum code length and restrict the code length of representative symbol to the representative symbol maximum code length. By restricting the code length of the representative symbol to the representative symbol maximum code length, code lengths of symbols represented by the representative symbol (i.e., code lengths of merge symbols) can be restricted to the maximum code length.
42 FIG. 41 FIG. 326 704 326 is a flowchart showing an example of the procedure of the second code length restriction process executed in the maximum code length restriction unit. The second code length restriction process corresponds to step Sof the maximum code length restriction process described above with reference to. The maximum code length restriction unitexecutes the second code length change process when there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length.
326 801 326 802 326 803 First, the maximum code length restriction unitselects an intermediate node whose depth from the root node is equal to the representative symbol maximum code length (step S). The maximum code length restriction unitidentifies non-representative symbols assigned to all leaf nodes (transitive leaf nodes) that can be reached by tracing from the selected intermediate node in child node direction (step S). The maximum code length restriction unitselects one non-representative symbol from the identified non-representative symbols (step S).
326 804 326 803 805 326 806 The maximum code length restriction unitchanges the code length of the selected non-representative symbol to the code length of the representative symbol (step S). The maximum code length restriction unitchanges the code lengths of the remaining non-representative symbols obtained by excluding the non-representative symbol selected in step Sfrom all the non-representative symbols assigned to the transitive leaf nodes, to the maximum code length (step S). Then, the maximum code length restriction unitchanges the code length of the representative symbol to the representative symbol maximum code length (step S), and ends the second code length change process.
326 With the above second code length change process, the maximum code length restriction unitcan restrict the code length of the representative symbol to the representative symbol maximum code length even when there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length.
802 42 FIG. When identifying the non-representative symbols assigned to the transitive leaf nodes in step Sof the second code length change process shown inis executed for a certain Huffman tree, the calculation amount may be increased. However, in a case where a specific combination of the maximum code length, the number of non-representative symbols, a code length of a representative symbol, and the maximum value of additional code lengths is used, patterns of a structure of a Huffman tree is limited, and thus the non-representative symbols assigned to the transitive leaf nodes can be identified with a small calculation amount.
43 FIG. 46 FIG. 43 FIG. 46 FIG. An example in which patterns of a structure of a Huffman tree are limited in a case where a specific combination of the maximum code length, the number of non-representative symbols, a code length of a representative symbol, and the maximum value of additional code lengths is used, will be described with reference toto. Here, it is assumed that the maximum code length is three, the number of non-representative symbols is eight, a code length of a representative symbol is three, and the maximum value of additional code lengths is two. The maximum code length is generally fixed to a value specified in a compression standard or the like. In addition, the number of non-representative symbols is a fixed value in a case where the Huffman tree is generated by using one or more representative symbols. Therefore, it is reasonable to assume that the maximum code length and the number of non-representative symbols are fixed. That is, into, under this assumption, a case where the code length of the representative symbol is three and the maximum value of additional code lengths is two is assumed. In this case, the representative symbol maximum code length is one (=3 −2) that is obtained by subtracting the maximum value of additional code lengths from the maximum code length. Therefore, the code length of the representative symbol is longer than the representative symbol maximum code length (3 >1) and violates a restriction on the representative symbol maximum code length. In other words, code lengths of merge symbols represented by the representative symbol violate a restriction on the maximum code length.
43 FIG. 46 FIG. In such a case, it will be considered that after clipping code lengths of the non-representative symbols, there is no non-representative symbol that can be swapped for the representative symbol and thus the violation of the restriction on the representative symbol maximum code length is not resolved. The non-representative symbol that can be swapped for the representative symbol is a non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length. That is, in this case, there is no non-representative symbol whose code length is shorter than or equal to the representative symbol maximum code length (i.e., there is no non-representative symbol whose code length is one). Furthermore, since the maximum code length is three, the minimum value of the code lengths of the non-representative symbols is either two or three. Therefore, the structure of the Huffman tree assumed in this case is limited to four patterns shown into.
43 FIG. shows a first pattern of the Huffman tree assumed in the above-described case. The first pattern corresponds to a case where the minimum value of the code lengths of the non-representative symbols is three.
90 912 911 921 391 43 FIG. 43 FIG. In this case, by selecting any four non-representative symbols from the non-representative symbols indicated in the Huffman tree, non-representative symbols corresponding to all of leaf nodes that can be reached by tracing from a selected intermediate node (for example, intermediate node) in a direction of child nodes can be selected. In the example shown in, the non-representative symbols “d”, “e”, “f”, and “g” are selected. Note that intermediate nodes that are transitive parents of the representative symbol “R” (in, intermediate nodesand) are excluded from target intermediate nodes from which the intermediate node selection unitselects an intermediate node.
44 FIG. shows a second pattern of the Huffman tree assumed in the above-described case. The second pattern corresponds to a case where the minimum value of the code lengths of the non-representative symbols is two and there is one non-representative symbol whose code length is two.
90 44 FIG. In this case, by selecting the one non-representative symbol whose code length is two and any two non-representative symbols whose code lengths are three from the non-representative symbols indicated by the Huffman tree, non-representative symbols corresponding to all of leaf nodes that can be reached by tracing from a selected intermediate node in a direction of child nodes can be selected. In the example shown in, the non-representative symbols “d”, “e”, and “f” are selected.
45 FIG. shows a third pattern of the Huffman tree assumed in the above-described case. The third pattern corresponds to a case where the minimum value of the code lengths of the non-representative symbols is two and there are two non-representative symbols whose code lengths are two.
90 45 FIG. In this case, by selecting the two non-representative symbols whose code lengths are two from the non-representative symbols indicated in the Huffman tree, non-representative symbols corresponding to all of leaf nodes that can be reached by tracing from a selected intermediate node in a direction of child nodes can be selected. In the example shown in, the non-representative symbols “d” and “f” are selected.
46 FIG. shows a fourth pattern of the Huffman tree assumed in the above-described case. The fourth pattern corresponds to a case where the minimum value of the code lengths of the non-representative symbols is two and there are two non-representative symbols whose code lengths are two but is different from the third pattern.
46 FIG. In the example shown in, the non-representative symbols “d” and “f” are also selected. In the fourth pattern, the nodes to which the non-representative symbols “b” and “c” are assigned are different from those in the third pattern.
326 326 326 Thus, the maximum code length restriction unitof the third embodiment can identify non-representative symbols corresponding to all of leaf nodes that can be reached by tracing from a selected intermediate node in a direction of child nodes, with a small calculation amount, in the case of using a specific combination of the maximum code length, the number of non-representative symbols, a code length of a representative symbol, and the maximum value of the additional code lengths. Note that, as described above, even when an example in which a structure of a Huffman tree is shown and symbols are assigned to leaf nodes is illustrated, the maximum code length restriction unitmay not actually manage the structure of the Huffman tree and the relationship between each leaf node and each symbol. For example, when managing a <representative symbol, code length> pair and a list of <non-representative symbol, code length> pairs, the maximum code length restriction unitcan perform the selection of symbols and the change of code lengths, and the like.
321 324 325 326 328 33 As described above, according to the first to third embodiments, processing time and processing resources can be reduced when code words that have respective code lengths shorter than or equal to an upper limit and are perfect codes are assigned to symbols, respectively. The frequency table generation unitgenerates a frequency table including N symbols, and N frequencies of occurrence that are associated with the N symbols, respectively, based on frequencies of occurrence of input symbols for each symbol. The Huffman tree generation unitgenerates a Huffman tree based on the frequency table. The code length determination unitdetermines N code lengths that correspond to the N symbols, respectively, based on the Huffman tree. In a case where the N code lengths include a first code length that is longer than the maximum code length, the maximum code length restriction unitselects a first symbol corresponding to the first code length from the N symbols, selects, from the N symbols, a second symbol corresponding to a second code length that is shorter than the maximum code length, changes the second code length corresponding to the second symbol to a code length that is obtained by adding one to the second code length, and changes the first code length corresponding to the first symbol to a code length that is equal to the changed second code length. The code assignment unitdetermines N variable length codes that are assigned to the N symbols, respectively, based on the N code lengths. The variable length coding unitconverts each of the input symbols into a variable length code, based on the N variable length codes that are assigned to the N symbols, respectively. N is an integer of two or more. The variable length code into which each of the input symbols is converted has a bit length between one bit length and the maximum code length inclusive.
15 15 32 15 32 With the configuration, the variable length coding devicecan proceed the maximum code length restriction process while maintaining variable length codes to be assigned as perfect codes. Accordingly, the variable length coding devicecan guarantee that the variable length codes assigned based on the code length of each symbol are perfect codes when the maximum code length restriction process has been completed. Therefore, the code table generation unitof the variable length coding devicecan reduce the processing time and the processing resources as compared to, for example, the code table generation unitA of the comparative example, which additionally executes the merge process.
Each of the various functions described in the first to third embodiments may be realized by a circuit (e.g., processing circuit). An exemplary processing circuit may be a programmed processor such as a central processing unit (CPU). The processor executes computer programs (instructions) stored in a memory thereby performs the described functions. The processor may be a microprocessor including an electric circuit. An exemplary processing circuit may be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a microcontroller, a controller, or other electric circuit components. The components other than the CPU described according to the embodiments may be realized in a processing circuit.
While certain embodiments have been described, these embodiments have been presented by way of example only, and are not intended to limit the scope of the inventions. Indeed, the novel devices and methods described herein may be embodied in a variety of other forms; furthermore, various omissions, substitutions and changes in the form of the embodiments described herein may be made without departing from the spirit of the inventions. The accompanying claims and their equivalents are intended to cover such forms or modification as would fall within the scope and spirit of the inventions.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
August 13, 2024
June 16, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.