Patentable/Patents/US-20260268068-A1
US-20260268068-A1

Information Processing Apparatus, Information Processing Method, and Computer Program

PublishedSeptember 10, 2026
Assigneenot available in USPTO data we have
Technical Abstract

For example, an information processing apparatus that automatically adds a new tag to content of a new genre is provided. An information processing apparatus includes: an existing tag proposal unit that presents, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal unit that generates a new tag appropriate for content to which a tag is to be added on the basis of a foundation model and presents the new tag to the annotator in a case where the annotator does not select the existing tag presented by the existing tag proposal unit; and a related content confirmation unit that presents content related to the new tag to the annotator for confirmation.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

An information processing apparatus comprising a new tag proposal unit that generates a new tag appropriate for content to which a tag is to be added on a basis of a foundation model and presents the new tag to an annotator.

2

claim 1 an existing tag proposal unit that presents, to the annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags, wherein the new tag proposal unit generates and presents a new tag in a case where the annotator does not select the existing tag presented by the existing tag proposal unit. . The information processing apparatus according to, further comprising

3

claim 2 the existing tag proposal unit estimates an existing tag appropriate for content to which a tag is to be added, using a learned model learned by using a sample with a correct answer label. . The information processing apparatus according to, wherein

4

claim 2 the new tag proposal unit generates a plurality of candidates for a new tag, and selects and presents an expression that is semantically far from the tag presented by the existing tag proposal unit from among the plurality of candidates. . The information processing apparatus according to, wherein

5

claim 1 a related content confirmation unit that presents content related to the tag to the annotator for confirmation. . The information processing apparatus according to, further comprising

6

claim 5 when the annotator selects the new tag presented by the new tag proposal unit, the related content confirmation unit presents content related to the new tag to the annotator for confirmation. . The information processing apparatus according to, wherein

7

claim 6 the new tag is added to content selected by the annotator from among the content presented by the related content confirmation unit. . The information processing apparatus according to, wherein

8

claim 7 a model is learned from a plurality of pieces of the content to which the new tag is added so that the tag can be added as an existing tag. . The information processing apparatus according to, wherein

9

an existing tag proposal step of presenting, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal step of generating, in a case where the annotator does not select the existing tag presented in the existing tag proposal step, a new tag appropriate for content to which a tag is to be added on a basis of a foundation model, and presenting the new tag to the annotator; and a related content confirmation step of presenting content related to the new tag to the annotator for confirmation when the annotator selects the new tag presented in the new tag proposal step. . An information processing method comprising:

10

an existing tag proposal unit that presents, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal unit that generates, in a case where the annotator does not select the existing tag presented by the existing tag proposal unit, a new tag appropriate for content to which a tag is to be added on a basis of a foundation model, and presents the new tag to the annotator; and a related content confirmation unit that presents content related to the new tag to the annotator for confirmation. . A computer program described in a computer readable format, the computer program causing a computer to function as:

Detailed Description

Complete technical specification and implementation details from the patent document.

The technology disclosed in the present specification (hereinafter, referred to as “present disclosure”) relates to an information processing apparatus, an information processing method, and a computer program that perform processing related to content management.

In order to effectively utilize a large amount of content such as music, tags are often added as metadata of the content. Although an annotator manually tags the content, in order to reduce the workload of the annotator, a method of automatically adding a tag using machine learning or the like is considered.

For example, there has been proposed a metadata generation system that automatically generates metadata of video content on the basis of a character code obtained by converting characters constituting a subtitle image displayed in superposition with video content such as television broadcasting by subtitle configuration character recognition and text information obtained by recognizing a voice included in the video content (see Patent Document 1).

An object of the present disclosure is to provide an information processing apparatus, an information processing method, and a computer program that automatically tag content.

The present disclosure has been made in view of the above problems, and a first aspect thereof is

an information processing apparatus including a new tag proposal unit that generates a new tag appropriate for content to which a tag is to be added on the basis of a foundation model and presents the new tag to an annotator.

The information processing apparatus according to the first aspect further includes an existing tag proposal unit that presents, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags. Then, the new tag proposal unit generates and presents a new tag in a case where the annotator does not select the existing tag presented by the existing tag proposal unit.

Furthermore, the information processing apparatus according to the first aspect further includes a related content confirmation unit that presents content related to the tag to the annotator for confirmation. When the annotator selects the new tag presented by the new tag proposal unit, the related content confirmation unit presents content related to the new tag to the annotator for confirmation.

Furthermore, a second aspect of the present disclosure is

an existing tag proposal step of presenting, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal step of generating, in a case where the annotator does not select the existing tag presented in the existing tag proposal step, a new tag appropriate for content to which a tag is to be added on the basis of a foundation model, and presenting the new tag to the annotator; and a related content confirmation step of presenting content related to the new tag to the annotator for confirmation when the annotator selects the new tag presented in the new tag proposal step. an information processing method including:

Furthermore, a third aspect of the present disclosure is

an existing tag proposal unit that presents, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal unit that generates, in a case where the annotator does not select the existing tag presented by the existing tag proposal unit, a new tag appropriate for content to which a tag is to be added on the basis of a foundation model, and presents the new tag to the annotator; and a related content confirmation unit that presents content related to the new tag to the annotator for confirmation. a computer program described in a computer readable format, the computer program causing a computer to function as:

The computer program according to the third aspect of the present disclosure is obtained by defining a computer program described in a computer readable format in such a way as to implement predetermined processing on a computer. The computer program can be provided for a computer capable of executing various program codes by a storage medium provided in a computer-readable form, a communication medium, a storage medium such as an optical disk, a magnetic disk, or a semiconductor memory, for example, or a communication medium such as a network. Then, by installing the computer program according to the third aspect of the present disclosure in a computer via any of the media, a cooperative action is exerted on the computer, and operational effects similar to those of the information processing apparatus according to the first aspect of the present disclosure can be achieved.

According to the present disclosure, it is possible to provide an information processing apparatus, an information processing method, and a computer program that automatically add a new tag to content.

Note that the effects described in the present specification are merely examples, and the effects to be brought by the present disclosure are not limited thereto. Furthermore, there are cases where the present disclosure further provides some other effects, in addition to the effects described above.

Still other objects, features, and advantages of the present disclosure will become apparent from a more detailed description based on embodiments as described later and the accompanying drawings.

A. Outline B. Foundation Model C. Basic Configuration D. Processing Procedure 1 D-. Processing Example (1) 2 D-. Processing Example (2) E. Annotation Screen Examples 1 E-: Annotation Screen Example (1) 2 E-: Annotation Screen Example (2) F. Configuration of Information Processing Apparatus Hereinafter, embodiments of the present disclosure will be described with reference to the drawings in the following order.

For example, automatic tagging of content can be realized by using a model subjected to machine learning to estimate a tag from content. If the number of samples for learning is large, machine learning for more appropriate tagging can be performed. Conversely, automation of tagging by machine learning is difficult with a small number of samples.

For example, in the music industry, music of a new genre is produced one after another, but in the case of content of a new genre, since the number of samples is small, it is difficult to automate tagging by machine learning. Furthermore, since the tag is abstract in the first place, there is an individual difference in interpretation of the tag for each annotator in a case where the annotator manually performs tagging. The tags attached to the same music are different for each annotator, and effective use of content using the tags (content classification and management) becomes difficult. Furthermore, if the annotator freely adds tags, the number of tags increases, or there are a lot of similar tags, so that it becomes difficult for the user who searches and recommends content using the annotator itself and the tag to grasp the content from the tag.

In short, it is desired to automatically add a new tag that is not similar to an existing tag to content of a new genre, but it is difficult to perform machine learning from a small number of samples. When the number of samples is small, the accuracy of the model is not improved, so that it is also difficult to add a new tag to existing content.

Therefore, the present disclosure proposes a technology for automatically adding a new tag to content using a foundation model learned with a large-scale sample. Since the foundation model is used, learning can be performed so that an appropriate tag can be generated even with a small number of samples.

According to the present disclosure, it is possible to propose a tag to be added to content using the foundation model, and thus, it is possible to absorb individual differences in tag interpretation by the annotator. Furthermore, according to the present disclosure, it is possible to propose related content to an annotator and refine a tag to be added to the content on the basis of selection of the annotator for the content. At the same time, by increasing the number of content samples associated with the new tag, the new tag can be incorporated into the existing tag system to enrich the system.

Therefore, according to the present disclosure, it is possible to add a tag that hardly causes a difference in interpretation to content (particularly, content of a new genre) using the foundation model having language knowledge, and thus, it is possible to reduce a burden on the annotator. Furthermore, according to the present disclosure, it is possible to add sophisticated tags to content of new genres and existing genres, and thus, it is possible to manage various content with a smaller number of tags by suppressing the randomization of similar tags that may occur in a case where tagging is performed with the sensitivity of the annotator.

The present disclosure has one feature in using a siner that automatically generates a new and appropriate tag to be added to content. In this section B, the foundation model will be described.

In conventional machine learning represented by deep learning, a method of performing learning using a sample with a correct answer label according to an application is common. In such a case, it is necessary to prepare a certain amount of samples for learning. Recently, a method of obtaining a highly accurate model by preparing a model obtained by preparing a large number of unsupervised samples in advance and performing self-supervised learning as a preliminary learning model and then performing fine tuning according to an application has become mainstream.

The foundation model is a model in which the direction of the latter method is further advanced, a huge amount of unsupervised samples are prepared, and self-supervised learning is performed, and a general-purpose model is constructed from large-scale data, and the general-purpose model is further customized according to the application.

One of the most famous examples of the foundation model is GPT-3 developed and published by OpenAI. GPT-3 is a model that learns 175 billion parameters using a large amount of unsupervised text data (45TB). The GPT-3 can be used for various applications of natural language such as generation, summarization, question response, and translation of sentences by devising a utilization method thereof. For example, a method has been studied in which necessary information is devised in the form of a prompt, provided to a foundation model, and information generated to solve a problem is exemplified, thereby appropriately solving the problem without changing parameters of the model itself.

The GPT-3 described above is an example of a foundation model specialized in text processing, but there is also a foundation model in which learning is performed with a very large number of samples by combining image information, voice or music information, and a relationship between these and text, and research and development of a foundation model for generating an image and a sound from text is also energetically performed.

For example, DALLE-2 is a foundation model developed and published by OpenAI to generate an image from text. Furthermore, AudioGen is a foundation model for generating sound from text. It is considered that these foundation models potentially hold not only text character strings but also relationships between image feature amounts and sound feature amounts in a huge parameter space through learning using a huge amount of samples. Therefore, the foundation model can utilize such relationships to generate bidirectionally between text and images and between text and sound.

1 FIG. 100 100 schematically illustrates a basic configuration of a tag adding systemthat adds a tag to content. The tag adding systemmay be configured as a part of a content editing system that performs various processing related to content editing, for example.

101 101 The terminalis a device that performs input/output operations for an annotator to add a tag to content, and includes a display and a console such as a keyboard, a mouse, and a touch panel. Furthermore, it is assumed that the terminalalso has a function of playing content to which a tag is to be added.

102 102 102 The content holding unitis a database that holds content to be tagged, such as text, audio, and image. For example, in a case where new content is created/edited by a content editing system (not illustrated), the new content is appropriately added to the content holding unit. For example, in the music content, a new song is appropriately added to the content holding unit.

103 103 102 The metadata holding unitis a database that holds a tag related to content as metadata. The metadata holding unitstores the tag added to the content held in the content holding unitin association with the content.

104 The foundation model unitholds foundation models related to text media and content media.

105 101 105 The existing tag proposal unitpresents an appropriate tag for content to which a tag is to be added from among existing (registered) tags via the terminal. The existing tag proposal unitestimates a tag appropriate for target content, for example, using a DNN model learned using a sample with a correct answer label. From the DNN model, a tag appropriate for the content is estimated from among existing (registered) tags learned as correct answer labels, and no new tag is generated.

104 106 101 106 Using the foundation model held in the foundation model unit, the new tag proposal unitgenerates a plurality of candidates for text information appropriate as a new tag related to content to which a tag is to be added, and presents the candidates to the annotator via the terminal. The new tag proposal unitgenerates a word or phrase appropriate to the tag as text information.

106 107 101 107 When the annotator adopts the new tag proposed by the new tag proposal unit, the related content confirmation unitpresents related content via the terminaland requests the annotator to confirm the content. This is because the newly generated tag corresponds to not only one target content but also other related content, and may be suitable for a tag of other content. The related content confirmation unitpresents a plurality of pieces of content related to the new tag to the annotator, and causes the annotator to select which content's tag is appropriate. As a result, it is possible to comprehensively associate the content corresponding to the new tag.

102 103 104 105 106 107 101 101 102 103 104 105 106 107 The content holding unit, the metadata holding unit, the foundation model unit, the existing tag proposal unit, the new tag proposal unit, and the related content confirmation unitmay be arranged in a cloud server, and may provide an automatic tagging service for content to the terminalas a client. Alternatively, the terminal, the content holding unit, the metadata holding unit, the foundation model unit, the existing tag proposal unit, the new tag proposal unit, and the related content confirmation unitmay all be arranged in a single device.

100 Next, a processing operation for adding a tag to content in the tag adding systemwill be described.

2 FIG. 100 illustrates a processing procedure for adding a tag to content on the tag adding systemin the form of a flowchart.

102 201 101 202 101 101 201 In a case where there is content to which a tag is to be added among the content held in the content holding unit(Yes in step S), the annotator selects the content and plays back the content on the terminal(step S). In the case of text or image content, the content is displayed on the display of the terminal, and in the case of sound content, the content is played back by the speaker of the terminal. Note that, in a case where there is no content to which a tag is to be added (No in step S), this processing ends.

201 102 102 In step S, the annotator selects, for example, new content added to the content holding unitas content to which a tag is to be added. For example, in the music content, a new song added to the content holding unitis selected.

105 202 101 203 105 Next, the existing tag proposal unitpresents one or a plurality of tags appropriate for the content selected in step Sfrom among the registered tags via the terminal(step S). The existing tag proposal unitestimates a tag appropriate for the content using the DNN model learned using the sample with a correct answer label.

202 105 203 204 101 103 102 204 205 In a case where an appropriate tag for the content selected in step Sis found from among the existing tags presented by the existing tag proposal unitin step S(Yes in step S), the annotator selects the tag via the terminal. In this case, the metadata holding unitstores the tag in association with the corresponding content in the content holding unit, assuming that the tag selected in step Sis added to the content (step S).

201 102 Thereafter, the processing returns to step S, and the above-described processing is repeatedly executed until there is no content to which a tag is to be added from the content holding unit(alternatively, until tag addition is completed for all the held content).

202 105 203 204 106 104 206 106 105 203 101 207 On the other hand, in a case where an appropriate tag for the content selected in step Shas not been found from among the tags presented by the existing tag proposal unitin step S(No in step S), subsequently, the new tag proposal unitgenerates a plurality of candidates for text information (word, phrase) appropriate as a new tag related to the content to which a tag is to be added, using the foundation model held in the foundation model unit(step S). Then, the new tag proposal unitselects an expression that is semantically far from the existing tag presented by the existing tag proposal unitin step Sfrom the plurality of generated candidates, and presents the expression to the annotator via the terminal(step S).

202 106 207 208 201 102 In a case where the annotator has not been able to find an appropriate tag for the content selected in step Sfrom among the tags presented by the new tag proposal unitin step S(No in step S), the annotator gives up tagging of this content. Then, the processing returns to step S, and the above-described processing is repeatedly executed until there is no content to which a tag is to be added from the content holding unit(alternatively, until tag addition is completed for all the held content).

202 106 207 208 101 On the other hand, in a case where an appropriate tag for the content selected in step Sis found from among the tags presented by the new tag proposal unitin step S(Yes in step S), the annotator selects the new tag via the terminal.

106 107 208 101 209 When the annotator selects the new tag proposed by the new tag proposal unit, there is a possibility that the new tag may correspond to other content. This is because the newly generated tag corresponds to not only one target content but also other related content, and may be suitable for a tag of other content. Therefore, the related content confirmation unitpresents the content related to the new tag selected in step Svia the terminal(step S), and requests the annotator to confirm the content.

208 210 103 208 210 211 The annotator selects content that is more appropriate for the new tag selected in step Sfrom among the plurality of pieces of related content presented (step S). Then, the metadata holding unitstores the new tag selected by the annotator in step Sin association with the content selected in step S(step S).

107 In this manner, the related content confirmation unitpresents a plurality of pieces of content related to the new tag to the annotator, and causes the annotator to select whether the newly generated tag is appropriate for any content, whereby the new tag can be comprehensively associated with the content. Moreover, regarding a content sample to be associated with a tag, the DNN model can be learned with the association as a correct label, and added as an existing tag.

201 102 Thereafter, the processing returns to step S, and the above-described processing is repeatedly executed until there is no content to which a tag is to be added from the content holding unit(alternatively, until tag addition is completed for all the held content).

207 2 FIG. Note that, in step Sin the flowchart illustrated in, processing of selecting a tag having a semantically far expression is performed, but a method of selecting a tag having a semantically far expression will be supplemented.

207 In natural language processing, a method of converting a language expression (word, phrase, sentence, document) that is a symbol sequence into a vector expression and defining a distance between the vectors is generally known. Also in step Sdescribed above, each tag can be converted into a vector representation, and the perspective of the semantic representation between the tags can be determined on the basis of the distance between the vectors.

(1) Bag of Words (2) Latent Semantic Indexing, Latent Direchlet Allocation (3) Word2vec As a method of vectorizing the language representation, for example, the following methods have been proposed so far, and these methods may be used, or vectorization (embedding) can be performed using a part of the internal representation of the foundation model.

Furthermore, as a distance definition between vectors, for example, a Euclidean distance and a cosine similarity can be exemplified. In the present embodiment, after each tag is converted into a vector representation, the perspective of the distance is quantified on the basis of the Euclidean distance or the cosine similarity between the vectors, whereby a tag that becomes a semantically far expression can be selected.

2 FIG. In the flowchart illustrated in, a processing procedure is performed in which a new tag is added in a case where there is no existing tag appropriate for content to which a tag is to be added. However, the existing tag and the new tag can be presented in parallel as tag candidates.

3 FIG. illustrates a processing procedure in a case where an existing tag and a new tag are added in parallel in the form of a flowchart.

102 301 101 302 101 101 301 In a case where there is content to which a tag is to be added among the content held in the content holding unit(Yes in step S), the annotator selects the content and plays back the content on the terminal(step S). In the case of text or image content, the content is displayed on the display of the terminal, and in the case of sound content, the content is played back by the speaker of the terminal. Note that, in a case where there is no content to which a tag is to be added (No in step S), this processing ends.

301 102 102 In step S, the annotator selects, for example, new content added to the content holding unitas content to which a tag is to be added. For example, in the music content, a new song added to the content holding unitis selected.

105 302 101 303 105 Next, the existing tag proposal unitpresents one or a plurality of tags appropriate for the content selected in step Sfrom among the registered tags via the terminal(step S). The existing tag proposal unitestimates a tag appropriate for the content to the annotator using the DNN model learned using the sample with a correct answer label.

106 104 304 106 105 303 101 305 305 106 Subsequently, the new tag proposal unitgenerates a plurality of candidates for text information (word, phrase) appropriate as a new tag related to content to which a tag is to be added, using the foundation model held in the foundation model unit(step S). Then, the new tag proposal unitselects an expression that is semantically far from the existing tag presented by the existing tag proposal unitin step Sfrom the plurality of generated candidates, and presents the expression to the annotator via the terminal(step S). In step S, after converting each tag into a vector expression, the new tag proposal unitcalculates an inter-vector distance such as Euclidean distance and cosine similarity, and selects a new tag that is an expression semantically far from the existing tag (same as above).

302 105 303 106 305 101 306 103 306 302 307 The annotator selects a tag suitable for the content selected in step Sfrom the existing tag presented by the existing tag proposal unitin step Sand the new tag presented by the new tag proposal unitin step Svia the terminal(step S). Then, the metadata holding unitstores the tag selected in step Sin association with the content selected in step S(step S).

301 102 Thereafter, the processing returns to step S, and the above-described processing is repeatedly executed until there is no content to which a tag is to be added from the content holding unit(alternatively, until tag addition is completed for all the held content).

4 FIG. 3 FIG. 4 FIG. 3 FIG. 106 illustrates a modification of the processing procedure illustrated inin the form of a flowchart. The processing procedure illustrated inis the same as the processing procedure illustrated inin that the existing tag and the new tag are added in parallel, but is different in that when the annotator selects the new tag proposed by the new tag proposal unit, processing of checking related content related to the new tag is added.

401 406 301 306 4 FIG. 3 FIG. Steps Sto Sin the flowchart illustrated inare common to steps Sto Sin the flowchart illustrated in, and thus description thereof is omitted here.

106 406 407 107 407 101 408 In a case where the annotator selects the new tag presented by the new tag proposal unitin step S(Yes in step S), there is a possibility that the new tag corresponds to other content (the same as above). Therefore, the related content confirmation unitpresents the content related to the new tag selected in step Svia the terminal(step S), and requests the annotator to confirm the content.

406 409 103 406 409 410 The annotator selects content that is more appropriate for the new tag selected in step Sfrom among the plurality of pieces of related content presented (step S). Then, the metadata holding unitstores the new tag selected by the annotator in step Sin association with the content selected in step S(step S).

107 In this manner, the related content confirmation unitpresents a plurality of pieces of content related to the new tag to the annotator, and causes the annotator to select whether the newly generated tag is appropriate for any content, so that it is possible to comprehensively associate the content corresponding to the new tag.

105 406 407 103 406 409 411 On the other hand, in a case where the annotator selects only the existing tag presented by the existing tag proposal unitin step S(No in step S), the metadata holding unitstores the existing tag selected by the annotator in step Sin association with the content selected in step S(step S).

401 102 Thereafter, the processing returns to step S, and the above-described processing is repeatedly executed until there is no content to which a tag is to be added from the content holding unit(alternatively, until tag addition is completed for all the held content).

101 In this section E, a configuration example of an annotation screen displayed on the display, which is used when the annotator performs an operation of adding a tag to content on the terminal, will be described. However, in the following, a case where music content is set as content to which a tag is to be added will be described.

2 FIG. 5 8 FIGS.to First, according to the flowchart illustrated in, annotation screens in the case of tagging content with an existing tag and a new tag in this order will be described with reference to.

5 FIG. 501 501 102 502 503 501 102 504 505 506 501 On the annotation screen illustrated in, the annotator inputs a song title of a music content to which a tag is to be added in a song title (Track) field. In a case where the song title input in the song title fieldis hit with the music content held in the content holding unit, the singer name and the lyrics of the music content are displayed in a singer name (Artist) fieldand a lyrics (Lyrics) field, respectively. Note that, in the song title (Track) field, a desired track name may be input as text, or each song title of the music content held in the content holding unitmay be displayed in a pull-down menu. Then, the annotator can actually listen to and confirm the music content by performing a playback operation of the selected music content using a playback button, a fast-forward button, and a rewind buttonimmediately below the song title field.

105 501 601 105 602 601 6 FIG. Next, the existing tag proposal unitestimates one or a plurality of existing tags appropriate for the music content selected in the song title field. Then, as illustrated in, on the annotation screen, a listof existing tags proposed by the existing tag proposal unitis displayed, and a new tag presentation (Show New Tag) buttonrequesting presentation of a new tag is further displayed. A check box is provided in each existing tag in the listof existing tags.

501 501 106 602 In a case where an appropriate existing tag is found for the music content selected in the song title field, the annotator can indicate tag addition to the music content by checking a check box. On the other hand, in a case where the annotator cannot find an appropriate existing tag for the music content selected in the song title field, the annotator can instruct the new tag proposal unitto present the new tag by pressing the new tag presentation (Show New Tag) button.

106 501 104 601 105 701 601 701 7 FIG. In response to an instruction from the annotator, the new tag proposal unitgenerates a plurality of candidates of text information (word, phrase) appropriate for the music content selected in the song title fieldusing the foundation model held in the foundation model unit, and selects an expression that is semantically far from the existing tagpresented by the existing tag proposal unit. Then, as illustrated in, a listof new tags that are semantically far from the existing tagis displayed on the annotation screen. A check box is provided in each new tag in the listof new tags.

501 701 107 801 801 7 FIG. 8 FIG. In a case where an appropriate new tag is found for the music content selected in the song title field, the annotator can indicate tag addition to the music content by checking a check box. When any new tag in the listof new tags is selected on the annotation screen illustrated in, as illustrated in, the related content confirmation unitfurther displays a pop-up windowlisting titles of content related to the selected new tag, and requests the annotator to confirm the content. This is because the newly generated tag corresponds to not only one target content but also other related content, and may be suitable for a tag of other content. At this time, it is more preferable to be able to play back music and display lyrics so that the content of the content can be confirmed by clicking the title of the displayed content. A check box is provided in the title of each content listed in the pop-up window.

801 103 801 In a case where the title of the content considered to be more appropriate for the selected new tag is found from the pop-up window, the annotator checks a check box and selects the content. Then, the metadata holding unitstores the new tag selected by the annotator in association with the content selected from the pop-up window.

3 FIG. 9 11 FIGS.to Next, following the flowchart illustrated in, annotation screens in a case where the existing tag and the new tag are tagged to content in parallel will be described with reference to.

9 FIG. 901 901 102 902 903 On the annotation screen illustrated in, the annotator inputs a song title of a music content to which a tag is to be added in a song title (Track) field. In a case where the song title input in the song title fieldis hit with the music content held in the content holding unit, the singer name and the lyrics of the music content are displayed in a singer name (Artist) fieldand a lyrics (Lyrics) field, respectively (same as above).

105 901 106 106 104 105 1001 105 1002 106 1001 1002 901 10 FIG. Next, the existing tag proposal unitestimates one or a plurality of existing tags appropriate for the music content selected in the song title field. Furthermore, the new tag proposal unitthe new tag proposal unitselects, from among the plurality of tag candidates generated using the foundation model held in the foundation model unit, a tag that is semantically far from the existing tag estimated by the existing tag proposal unit. Then, as illustrated in, on the annotation screen, a listof existing tags proposed by the existing tag proposal unitand a listof new tags proposed by the new tag proposal unitare displayed in parallel. A check box is provided in each tag of the listof existing tags and the listof new tags. In a case where the appropriate tag is found for the musical content selected in the song title field, the annotator can indicate the tag addition to the musical content by checking the check box.

1002 107 1101 1101 11 FIG. When any new tag in the listof new tags is selected, as illustrated in, the related content confirmation unitfurther displays a pop-up windowlisting titles of content related to the selected new tag, and requests the annotator to confirm the content. This is because the newly generated tag corresponds to not only one target content but also other related content, and may be suitable for a tag of other content. A check box is provided in the title of each content listed in the pop-up window.

1101 103 1101 In a case where the title of the content considered to be more appropriate for the selected new tag is found from the pop-up window, the annotator checks a check box and selects the content. Then, the metadata holding unitstores the new tag selected by the annotator in association with the content selected from the pop-up window.

Expressions such as “broken heart” and “lost love” can be considered as a tag indicating a love break for a music content singing a love break. If the annotator is allowed to freely add a tag, there is a concern that a plurality of tags representing a similar concept is set particularly in a case where a plurality of annotators performs work, and the entire metadata becomes unclear. By adding a new tag generated using the foundation model to the music content as in the present disclosure, it is possible to suppress such disorder of similar tags, and as a result, it is possible to manage various contents with a smaller number of tags as metadata.

12 FIG. 2000 2000 1000 2000 105 106 107 2000 101 In this section F, an information processing apparatus that is used in implementing the present disclosure is described.illustrates a configuration example of the information processing apparatus. The information processing apparatuscan constitute the entire tag adding systemor a part thereof. The information processingcan constitute, for example, any one or two or more of the existing tag proposal unit, the new tag proposal unit, and the related content confirmation unit. Furthermore, the information processing apparatuscan also constitute the terminaloperated by the annotator.

2000 2001 2002 2003 2004 2005 2006 2007 2008 2009 2010 2011 2013 12 FIG. The information processing apparatusillustrated inincludes a central processing unit (CPU), a read only memory (ROM), a random access memory (RAM), a host bus, a bridge, an expansion bus, an interface unit, an input unit, an output unit, a storage unit, a drive, and a communication unit.

2001 2000 2002 2001 2003 2001 2003 2001 The CPUcontrols the overall operation of the information processing apparatusaccording to various programs. The ROMstores programs (such as a basic input/output system) and computation parameters to be used by the CPUin a nonvolatile manner. The RAMis used to load a program to be used in execution of the CPUand temporarily store parameters such as work data that appropriately changes during program execution. Examples of the program to be loaded into the RAMand executed by the CPUinclude various application programs, an operating system (OS), and the like.

2001 2002 2003 2004 2001 2002 2003 2000 105 106 107 101 5 12 FIGS.to The CPU, the ROM, and the RAMare interconnected by the host busincluding a CPU bus or the like. Then, the CPUoperates in conjunction with the ROMand the RAMto execute various application programs under an execution environment provided by the OS, thereby enabling various functions and services to be implemented. In a case where the information processing apparatusis a personal computer, the OS is, for example, Windows (registered trademark), Unix (registered trademark) or the like of Microsoft Corporation in the US. Furthermore, the application program includes a program for operating as at least one of the existing tag proposal unit, the new tag proposal unit, or the related content confirmation unit. Furthermore, the application program may also include a program to operate as terminaloperated by the annotator (or displays the annotation screens illustrated in).

2004 2006 2005 2006 2005 2000 2004 2005 2006 The host busis connected to the expansion busvia the bridge. The expansion busis, for example, a peripheral component interconnect (PCI) bus or PCI Express, and the bridgeis based on the PCI standard. However, the information processing apparatusdoes not necessarily have a configuration in which circuit components are separated by the host bus, the bridge, and the expansion bus, and may be designed in such a manner that almost all circuit components are interconnected by a single bus (not illustrated).

2007 2008 2009 2010 2011 2013 2006 2000 2000 2000 9 FIG. The interface unitconnects peripheral devices such as the input unit, the output unit, the storage unit, the drive, and the communication unitaccording to the standard of the expansion bus. Note that all of the peripheral devices illustrated inare not necessarily essential, and the information processing apparatusmay further include a peripheral device not illustrated in the drawing. Furthermore, the peripheral device may be built in the main body of the information processing apparatus, or some peripheral devices may be externally connected to the main body of the information processing apparatus.

2008 2001 2000 2008 2009 The input unitincludes an input control circuit that generates an input signal on the basis of an input from a user to output the input signal to the CPU, and the like. In a case where the information processing apparatusis a personal computer, the input unitmay include a keyboard, a mouse, and a touch panel, and may further include a camera and a microphone. Furthermore, the output unitincludes a display device such as a liquid crystal display (LCD) device, an organic electro-luminescence (EL) display device, or a light emitting diode (LED), for example.

2010 2001 2010 2010 102 103 104 The storage unitstores files such as programs (application, OS, etc.) to be executed by the CPUand various pieces of data. Although the storage unitincludes, for example, a mass storage device such as a solid state drive (SSD) or a hard disk drive (HDD), it may include an external storage device. The storage unitoperates as, for example, at least one of the content holding unit, the metadata holding unit, or the foundation model unit.

2012 2011 113 2011 2012 2003 2010 2003 2010 2012 A removable recording mediumincludes a cartridge-type storage medium such as a micro-SD card, for example. The driveperforms reading and writing operations on a removable storage mediumloaded therein. The driveoutputs data read from the removable recording mediumto the RAMand the storage unit, and writes data on the RAMand the storage unitto the removable recording medium.

2013 2013 The communication unitis a device that performs wireless communication such as Wi-Fi (registered trademark), Bluetooth (registered trademark), or a cellular communication network such as 4G or 5G. Furthermore, the communication unitalso include a terminal such as a universal serial bus (USB) or a high-definition multimedia interface (HDMI: registered trademark), and may further include a function of performing HDMI (registered trademark) communication with a USB device such as a scanner or a printer, a display, or the like.

The present disclosure is described in detail with reference to the specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and those skilled in the art obviously can make modifications and substitutions of the embodiments without departing from the gist of the present disclosure. Furthermore, the effects described in the present specification are each merely an example, thus, the effects brought by an embodiment of the present disclosure are not limited and may include an additional effect that is not described herein.

In the present specification, the present disclosure has been described mainly with respect to an embodiment in which the tag is added to the music content, but the gist of the present disclosure is not limited thereto. Similarly, when a tag is added to content of various media such as moving image content, movie content, text content, or the like, the present disclosure is applied, so that a difference in interpretation is hardly generated with respect to the content, a new and sophisticated tag can be added, and a burden on the annotator can be reduced.

In short, the present disclosure is described in an illustrative manner, and the content disclosed in the present specification should not be interpreted in a limited manner. To determine the subject matter of the present disclosure, the claims should be taken into consideration.

The series of processing described in the present specification can be executed by hardware, software, or a configuration in which hardware and software are combined. In a case where the processing is executed by software, a program recorded with a processing sequence related to implementation of the present disclosure is installed and executed in a memory incorporated in dedicated hardware in a computer. It is also possible to install a program in a general-purpose computer capable of executing various types of processing and cause the computer to execute the processing related to implementation of the present disclosure.

The program can be preliminarily stored in a recording medium provided in the computer, such as an HDD, an SSD, or a ROM. Alternatively, the program can be temporarily or permanently stored in a removable recording medium such as a flexible disk, a compact disc read only memory (CD-ROM), a magneto optical (MO) disk, a digital versatile disc (DVD), a Blu-ray Disc (registered trademark) (BD), a magnetic disk, or a universal serial bus (USB) memory. Using such a removable recording medium enables a program related to implementation of the present disclosure as so-called package software to be provided.

Furthermore, the program may be transferred from a download site to a computer in a wireless or wired manner via a network such as a wide area network (WAN) typified by a cellular network, a local area network (LAN), or the Internet. The computer can receive the program thus transferred and cause the program to be installed in a mass storage device such as an HDD or an SSD in the computer.

(1) An information processing apparatus including a new tag proposal unit that generates a new tag appropriate for content to which a tag is to be added on the basis of a foundation model and presents the new tag to an annotator. (2) The information processing apparatus according to (1) described above, further including an existing tag proposal unit that presents, to the annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags, in which the new tag proposal unit generates and presents a new tag in a case where the annotator does not select the existing tag presented by the existing tag proposal unit. (3) The information processing apparatus according to (2) Note that the present disclosure may also have the following configurations.

the existing tag proposal unit estimates an existing tag appropriate for content to which a tag is to be added, using a learned model learned by using a sample with a correct answer label. (4) The information processing apparatus according to any one of (2) or (3) described above, in which the new tag proposal unit generates a plurality of candidates for a new tag, and selects and presents an expression that is semantically far from the tag presented by the existing tag proposal unit from among the plurality of candidates. (5) The information processing apparatus according to any one of (1) to (4) described above, further including a related content confirmation unit that presents content related to the tag to the annotator for confirmation. (6) The information processing apparatus according to (5) described above, in which when the annotator selects the new tag presented by the new tag proposal unit, the related content confirmation unit presents content related to the new tag to the annotator for confirmation. (7) The information processing apparatus according to (6) described above, in which the new tag is added to content selected by the annotator from among the content presented by the related content confirmation unit. (8) The information processing apparatus according to (7) described above, in which a model is learned from a plurality of pieces of the content to which the new tag is added so that the tag can be added as an existing tag. (9) An information processing method including: an existing tag proposal step of presenting, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal step of generating, in a case where the annotator does not select the existing tag presented in the existing tag proposal step, a new tag appropriate for content to which a tag is to be added on the basis of a foundation model, and presenting the new tag to the annotator; and a related content confirmation step of presenting content related to the new tag to the annotator for confirmation when the annotator selects the new tag presented in the new tag proposal step. (10) A computer program described in a computer readable format, the computer program causing a computer to function as: an existing tag proposal unit that presents, to an annotator, an existing tag appropriate for content to which a tag is to be added from among existing tags; a new tag proposal unit that generates, in a case where the annotator does not select the existing tag presented by the existing tag proposal unit, a new tag appropriate for content to which a tag is to be added on the basis of a foundation model, and presents the new tag to the annotator; and a related content confirmation unit that presents content related to the new tag to the annotator for confirmation. described above, in which

100 Tag adding system 101 Terminal 102 Content holding unit 103 Metadata holding unit 104 Foundation model unit 105 Existing tag proposal unit 106 New tag proposal unit 107 Related content confirmation unit 2000 Information processing apparatus 2001 CPU 2002 ROM 2003 RAM 2004 Host bus 2005 Bridge 2006 Expansion bus 2007 Interface unit 2008 Input unit 2009 Output unit 2010 Storage unit 2011 Drive 2012 Removable recording medium 2013 Communication unit

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

January 26, 2024

Publication Date

September 10, 2026

Inventors

YASUHARU ASANO

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND COMPUTER PROGRAM” (US-20260268068-A1). https://patentable.app/patents/US-20260268068-A1

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.

INFORMATION PROCESSING APPARATUS, INFORMATION PROCESSING METHOD, AND COMPUTER PROGRAM — YASUHARU ASANO | Patentable