A display apparatus and a voice controlling method thereof are provided. The voice controlling method includes receiving a voice of a user; converting the voice into text; and sequentially changing and applying a plurality of different determination criteria to the text until a control operation corresponding to the text is determined; and performing the determined control operation to control the display apparatus.
Legal claims defining the scope of protection, as filed with the USPTO.
receiving a voice of a user; converting the voice into text; determining a control operation corresponding to the text by sequentially applying a plurality of different determination criteria to the text; and performing the control operation to control the display apparatus, wherein the plurality of different determination criteria comprise criteria of whether the text corresponds to a title of an object displayed on a screen of the display apparatus, and whether the text corresponds to a stored command, a criterion of whether the text corresponds to a title of an object displayed on a screen of the display apparatus being applied before a criterion of whether the text corresponds to a stored command, and determining whether the text corresponds to the title of the object; in response to determining that the text corresponds to the title of the object, determining the control operation based on the object; in response to determining that the text does not correspond to the title of the object, determining whether the text corresponds to the stored command; and in response to determining that the text corresponds to the stored command, determining an operation corresponding to the stored command as the control operation. wherein the determining the control operation comprises: . A voice controlling method of a display apparatus, the voice controlling method comprising:
claim 1 determining an operation corresponding to the object as the control operation. . The voice controlling method of, wherein the determining the control operation based on the object comprises:
claim 1 . The voice controlling method of, wherein the determining whether the text corresponds to the title of the object comprises, in response to a part of the title of the object being displayed and the text corresponding to at least a portion of the displayed part of the object, determining that the text corresponds to the title of the object.
claim 1 . The voice controlling method of, wherein the determining whether the text corresponds to the title of the object comprises, in response to only a part of one word included in the title of the object being displayed and the text corresponding to the whole one word, determining that the text corresponds to the title of the object.
claim 1 . The voice controlling method of, wherein the object comprises at least one of a content title, an image title, a text icon, a menu name, and a number that are displayed on the screen.
claim 1 . The voice controlling method of, wherein the stored command comprises at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus.
claim 1 in response to determining that the text does not correspond to the stored command, determining whether a meaning of the text is analyzable; and in response to determining that the meaning of the text is analyzable, analyzing the meaning of the text and determining an operation of displaying a response message corresponding to a result of the analyzing as the control command. . The voice controlling method of, further comprising:
claim 7 in response to determining that the meaning of the text is not analyzable, determining an operation of a search using the text as a keyword, as the control operation. . The voice controlling method of, further comprising:
claim 1 . The method of, wherein the plurality of different determination criteria further comprise criteria of whether the text is grammatically analyzable, and whether the text refers to a keyword.
a voice input circuit configured to receive a voice of a user; a voice converter configured to convert the voice into text; a storage configured to store a plurality of determination criteria that are different from one another; and a controller configured to determine a control operation corresponding to the text by sequentially applying a plurality of different determination criteria to the text, and perform the control operation, wherein the plurality of different determination criteria comprise criteria of whether the text corresponds to a title of an object displayed on a screen of the display apparatus, and whether the text corresponds to a stored command, a criterion of whether the text corresponds to a title of an object displayed on a screen of the display apparatus being applied before a criterion of whether the text corresponds to a stored command, and wherein the controller is further configured to determine whether the text corresponds to the title of the object, determine the control operation based on the object in response to determining that the text corresponds to the title of the object, determine whether the text corresponds to the stored command in response to determining that the text does not correspond to the title of the object, and determine an operation corresponding to the stored command as the control operation in response to determining that the text corresponds to the stored command. . A display apparatus comprising:
claim 10 . The display apparatus of, wherein the controller is further configured to, in response to determining that the text corresponds to the title of the object, determine an operation corresponding to the object as the control operation.
claim 10 . The display apparatus of, wherein the controller is further configured to, in response to only a part of the title of the object being displayed and determining that the text corresponds to the part of the title of the object being displayed, determine that the text corresponds to the title of the object.
claim 10 . The display apparatus of, wherein the controller is further configured to, in response to only a part of one word included in the title of the object being displayed and determining that the text corresponds to the whole one word, determine that the text corresponds to the title of the object.
claim 10 . The display apparatus of, wherein the object comprises at least one of a content title, an image title, a text icon, a menu name, and a number that are displayed on the screen.
claim 10 . The display apparatus of, wherein the stored command comprises at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus.
claim 10 . The display apparatus of, wherein the controller is further configured to, in response to determining that the text does not correspond to the stored command, determine whether a meaning of the text is analyzable, and, in response to determining that the meaning of the text is analyzable, analyze the meaning of the text and determine an operation of displaying a response message corresponding to the analysis result, as the control operation.
claim 16 . The display apparatus of, wherein the controller is further configured to, in response to determining that the meaning of the text is not analyzable, determine an operation of a search using the text as a keyword, as the control operation.
a display; a voice receiving circuit; and control the display to display at least one title of content on a screen of the display; receive, via the voice receiving circuit, a voice input of a user while the at least one title of content is displayed on the screen of the display; obtain text resulting from processing the received voice input; based on the obtained text corresponding to a title of content among the at least one title of content displayed on the screen of the display, control the display to display content corresponding to the title of content displayed on the screen of the display, regardless of whether the obtained text matches any predefined instructions for controlling the display device; based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and matching an instruction among the predefined instructions for controlling the display device, execute the instruction, and control the display to display a user interface, the user interface indicating a result of executing the instruction while the at least one title of content is displayed on the screen of the display; and based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and not matching any of the predefined instructions for controlling the display device, control the display to display a search result regarding the obtained text, at least one processor configured to: wherein the at least one processor is further configured to identify whether the obtained text corresponds to the title of content before identifying whether the obtained text matches any of the predefined instructions. 18. A display device comprising:
claim 18 19. The display device of, wherein the at least one processor is further configured to, in response to the obtained text corresponding to the title of content among the at least one title of content displayed on the screen of the display, obtain a first control information for executing a first operation irrespective of whether the text corresponds to any of the predefined instructions for controlling the display device.
claim 19 in response to the obtained text corresponding to the title of content among the at least one title of content displayed on the screen of the display, obtain the first control information; and in response to the obtained text not corresponding to text associated with any of the at least one title of content displayed on the screen of the display and corresponding to the instruction among the predefined instructions for controlling the display device, obtain a second control information for executing a second operation. 20. The display device of, wherein the at least one processor is further configured to:
claim 20 wherein the at least one processor is further configured to: transmit the voice input to a first server connected to the display device; and receive text corresponding to the transmitted voice input from the first server. 21. The display device of, wherein the display device further comprises a communicator;
claim 21 22. The display device of, wherein the first control information and the second control information are obtained from a second server connected to the display device and the first server.
claim 22 based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and not matching any of the predefined instructions for controlling the display device, control the communicator to transmit information regarding the obtained text to a third server; receive the search result through the communicator from the third server; and control the display to display the search result received from the third server. 23. The display device of, wherein the at least one processor is further configured to:
claim 18 wherein the voice input of the user is received through the microphone. 24. The display device of, wherein the display device further comprises a microphone; and
claim 18 25. The display device of, wherein the voice input of the user is received from a remote control device for controlling the display device.
claim 18 26. The display device of, wherein the at least one processor is further configured to, based on the search result comprises a plurality of search results, control the display to display a plurality of numbers corresponding to each of the plurality of search results together with the plurality of search results.
displaying at least one title of content on a screen of the display; receiving via a voice receiving circuit, a voice input of a user while the at least one title of content is displayed on the screen of the display; obtaining text resulting from processing the received voice input; based on the obtained text corresponding to a title of content among the at least one title of content displayed on the screen of the display, displaying content corresponding to the title of content displayed on the screen of the display, regardless of whether the obtained text matches any of predefined instructions for controlling the display device; based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and matching an instruction among the predefined instructions for controlling the display device, executing the instruction, and displaying a user interface, the user interface indicating a result of executing the instruction while the at least one title of content is displayed on the screen of the display; and based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and not matching any of the predefined instructions for controlling the display device, displaying a search result regarding the obtained text, wherein identifying whether the obtained text corresponds to the title of content is performed before identifying whether the obtained text matches any of the predefined instructions. 27. A method of a display device, comprising:
claim 27 in response to the obtained text corresponding to the text associated with the title of content among the at least one title of content displayed on the screen of the display, obtaining a first control information for executing a first operation irrespective of whether the text corresponds to any of the predefined instructions for controlling the display device. 28. The method of, further comprising:
claim 28 in response to the obtained text corresponding to the text associated with the title of content among the at least one title of content displayed on the screen of the display, obtaining the first control information; and in response to the obtained text not corresponding to text associated with any of the at least one title of content displayed on the screen of the display and corresponding to the instruction among the predefined instructions for controlling the display device, obtaining a second control information for executing a second operation. 29. The method of, wherein the method further comprises:
claim 29 transmitting the voice input to a first server connected to the display device; and receiving text corresponding to the transmitted voice input from the first server. 30. The method of, wherein the method further comprises:
claim 30 31. The method of, wherein the first control information and the second control information are obtained from a second server connected to the display device and the first server.
claim 31 based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and not matching any of the predefined instructions for controlling the display device, transmitting information regarding the obtained text to a third server; receiving the search result from the third server; and displaying the search result received from the third server. 32. The method of, wherein the method further comprises:
claim 27 33. The method of, wherein the voice input of the user is received through a microphone of the display device.
claim 27 34. The method of, wherein the voice input of the user is received from a remote control device for controlling the display device.
claim 27 35. The method of, wherein the method further comprises, based on the search result comprises a plurality of search results, displaying a plurality of numbers corresponding to each of the plurality of search results together with the plurality of search results.
displaying at least one title of content on a screen of the display; receiving via a voice receiving circuit, a voice input of a user while the at least one title of content is displayed on the screen of the display; obtaining text resulting from processing the received voice input; based on the obtained text corresponding to a title of content among the at least one title of content displayed on the screen of the display, displaying content corresponding to the title of content displayed on the screen of the display, regardless of whether the obtained text matches any of predefined instructions for controlling the display device; based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and matching an instruction among the predefined instructions for controlling the display device, executing the instruction, and displaying a user interface, the user interface indicating a result of executing the instruction while the at least one title of content is displayed on the screen of the display; and based on the obtained text not corresponding to any of the at least one title of content displayed on the screen of the display and not matching any of the predefined instructions for controlling the display device, displaying a search result regarding the obtained text, wherein identifying whether the obtained text corresponds to the title of content is performed before identifying whether the obtained text matches any of the predefined instructions. 36. A computer readable medium which includes a program for executing a method of a display device, the method comprising:
Complete technical specification and implementation details from the patent document.
More than one reissue application has been filed for the reissue of U.S. Pat. No. 9,711,149. The reissue applications are (1) the present application, (2) U.S. patent application Ser. No. 17/175,261, filed on Feb. 12, 2021, and (3) U.S. patent application Ser. No. 16/515,466, filed on July 18, 2019. U.S. patent application Ser. No. 16/515,466, filed on Jul. 18, 2019, is a reissue of U.S. Pat. No. 9,711,149, which was filed as U.S. patent application Ser. No. 14/515,781 on Oct. 16, 2014 and claims priority from Korean Patent Application No. 10-2014-0009388, filed on Jan. 27, 2014, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety. U.S. patent application Ser. No. 17/175,261, filed on Feb. 12, 2021, is a continuation reissue of U.S. patent application Ser. No. 16/515,466. The present application is a continuation reissue of U.S. patent application Ser. No. 17/175,261.
This application claims priority from Korean Patent Application No. 10-2014-0009388, filed on Jan. 27, 2014, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety.
1. Field
Methods and devices of manufacture consistent with exemplary embodiments relate to a display apparatus and a voice controlling method thereof, and more particularly, to a display apparatus for determining a voice input of a user to perform an operation and a voice controlling method thereof.
2. Description of the Related Art
As display apparatuses have been gradually becoming more multifunctional and advanced, various input methods for controlling the display apparatuses have been developed. For example, an input method using a voice control technology, an input method using a mouse, an input method using a touch pad, an input method using a motion sensing remote controller, etc. have been developed.
However, there are several kinds of disadvantages in using voice control technology. For example, if a voice uttered by a user is a simple keyword having no verb, a different operation from that intended by the user may be performed.
In other words, if the display apparatus misrecognizes the voice uttered by the user, the display apparatus may not be controlled as the user wants.
Exemplary embodiments address at least the above disadvantages and other disadvantages not described above. Also, the exemplary embodiments are not required to overcome the disadvantages described above, and an exemplary embodiment may not overcome any of the disadvantages described above.
One or more exemplary embodiments provide a display apparatus for determining a voice input of a user to perform an operation corresponding to an intention of the user and a method of controlling a voice.
According to an aspect of an exemplary embodiment, there is provided a voice controlling method of a display apparatus, the method including: receiving a voice of a user; converting the voice into text; sequentially changing and applying a plurality of different determination criteria to the text until a control operation corresponding to the text is determined; and performing the determined control operation to control the display apparatus.
The sequentially changing and applying the plurality of different determination criteria may include determining whether the text corresponds to a title of an object displayed on a screen of the display apparatus; and in response to determining that the text corresponds to the title of the object, determining an operation corresponding to the object as the control operation.
The determining whether the text corresponds to the title of the object may include, in response to a part of the title of the object being displayed and the text corresponding to at least a portion of the displayed part of the object, determining that the text corresponds to the title of the object.
The determining whether the text corresponds to the title of the object may include, in response to only a part of one word included in the title of the object being displayed and the text corresponding to the whole one word, determining that the text corresponds to the title of the object.
The object may include at least one of a content title, an image title, a text icon, a menu name, and a number that are displayed on the screen.
The sequentially changing and applying the plurality of different determination criteria may further include: in response to determining that the text does not correspond to the title of the object, determining whether the text corresponds to a stored command; and in response to determining that the text corresponds to the stored command, determining an operation corresponding to the stored command as the control operation.
The stored command may include at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus.
The sequentially changing and applying the plurality of different determination criteria may include: in response to determining that the text does not correspond to the stored command, determining whether a meaning of the text is analyzable; and in response to determining that the meaning of the text is analyzable, analyzing the meaning of the text and determining an operation of displaying a response message corresponding to the analysis result as the control command.
The sequentially changing and applying the plurality of different determination criteria may include in response to determining that the meaning of the text is not analyzable, determining an operation of a search using the text as a keyword, as the control operation.
According to an aspect of another exemplary embodiment, there is provided a display apparatus including: a voice input circuit configured to receive a voice of a user; a voice converter configured to convert the voice into text; a storage configured to store a plurality of determination criteria that are different from one another; and a controller configured to sequentially change and apply a plurality of different determination criteria to the text until a control operation corresponding to the text is determined, and perform the determined control operation.
The controller may be configured to sequentially change and apply the plurality of different determination criteria to the text by determining whether the text corresponds to a title of an object displayed on a screen of the display apparatus and, in response to determining that the text corresponds to the title of the object, determining an operation corresponding to the object as the control operation.
The controller may be configured to, in response to only a part of the title of the object being displayed and determining that the text corresponds to the part of the title of the object being displayed, determine that the text corresponds to the title of the object.
The controller may be configured to, in response to only a part of one word included in the title of the object being displayed and determining that the text corresponds to the whole one word, determine that the text corresponds to the title of the object.
The object may include at least one of a content title, an image title, a text icon, a menu name, and a number that are displayed on the screen.
The controller may be configured to sequentially change and apply the plurality of different determination criteria to the text by, in response to determining that the text does not correspond to the title of the object, determining whether the text corresponds to a stored command, and in response to determining that the text corresponds to the stored command, determining an operation corresponding to the stored command as the control operation.
The stored command may include at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus.
The controller may be configured to sequentially change and apply the plurality of different determination criteria to the text by, in response to determining that the text does not correspond to the stored command, determining whether a meaning of the text is analyzable, and, in response to determining that the meaning of the text is analyzable, analyzing the meaning of the text and determines an operation of displaying a response message corresponding to the analysis result, as the control operation.
The controller may be configured to sequentially change and apply the plurality of different determination criteria to the text by, in response to determining that the meaning of the text is not analyzable, determining an operation of a search using the text as a keyword, as the control operation.
The plurality of different determination criteria may include criteria of whether the text corresponds to a title of a displayed object, whether the text corresponds to a stored command, whether the text is grammatically analyzable, and whether the text refers to a keyword.
According to an aspect of another exemplary embodiment, there may be provided a voice controlling method of a display apparatus, the voice controlling method including receiving a voice of a user; converting the voice into text; applying at least two tiers of hierarchical criteria the text to determine a control operation corresponding to the text; and controlling the display apparatus according to the determined control operation.
Exemplary embodiments are described in greater detail with reference to the accompanying drawings.
In the following description, the same drawing reference numerals are used for the same elements even in different drawings. The matters defined in the description, such as detailed construction and elements, are provided to assist in a comprehensive understanding of the exemplary embodiments. Thus, it is apparent that the exemplary embodiments can be carried out without those specifically defined matters. Also, well-known functions or constructions are not described in detail since they would obscure the exemplary embodiments with unnecessary detail.
1 FIG. 1 FIG. 100 110 120 140 130 is a block diagram illustrating a structure of a display apparatus according to an exemplary embodiment. Referring to, a display apparatusincludes a voice input circuit, a voice converter, a controller, and a storage.
100 110 120 100 The display apparatusmay receive a voice of a user through the voice input circuitand convert the voice into text using the voice converter. Here, the display apparatusmay sequentially change a plurality of different determination criteria until a control operation corresponding to the converted text is determined and then determine the control operation corresponding to the converted text.
100 100 The display apparatusmay be a display apparatus such as a smart TV, but this is only an . Alternatively, the display apparatusmay be realized as, for example, a desktop personal computer (PC), a tablet PC, a smartphone, or the like or may be realized as another type of input device such as a voice input device.
110 110 100 100 110 120 The voice input circuitis an element that receives the voice of the user. In detail, the voice input circuitmay include a microphone and associated circuitry to directly receive the user's voice as sound and convert the sounds to an electric signal, or may include circuitry to receive an electric signal corresponding to the user's voice input to the display apparatusthrough a microphone that is connected to the display apparatusby wire or wireless. The voice input circuittransmits the signal corresponding to the user's voice to the voice converter.
120 The voice converterparses a waveform of a characteristic of the user's voice signal (i.e., a characteristic vector of the user's voice signal) to recognize words or a word string corresponding to the voice uttered by a user and outputs the recognized words as text information.
120 120 120 In detail, the voice convertermay recognize the words or the word string uttered by the user from the user voice signal by using at least one of various recognition algorithms such as a dynamic time warping method, a Hidden Markov model, a neural network, etc. and convert the recognized voice into text. For example, if the Hidden Markov model is used, the voice converterrespectively models a time change and a spectrum change of the user voice signal to detect a similar word from a stored language database (DB). Therefore, the voice convertermay output the detected word or words as text.
110 120 100 110 120 The voice input circuitand the voice converterhave been described as elements that are installed in the display apparatusin the present exemplary embodiment, but this is only an example. Alternatively, the voice input circuitand the voice convertermay be realized as external devices.
140 110 140 140 110 120 140 130 140 100 The controllerperforms a control operation corresponding to the user voice input through the voice input circuit. The controllermay start a voice input mode according to a selection of the user. If the voice input mode starts, the controllermay activate the voice input circuitand the voice converterto receive the user's voice. If the user voice is input when a voice input mode is active, the controlleranalyzes an intention of the user by using a plurality of different determination criteria stored in the storage. The controllerdetermines a control operation according to the analysis result to perform an operation of the display apparatus
140 140 140 In detail, the controllerdetermines whether the converted text corresponds to a title of an object in a displayed screen. If the converted text corresponds to the title of the object, the controllerperforms an operation corresponding to the object. For example, the controllermay perform an operation matching the object. In detail, the object may include at least one of a content title, an image title, a text icon, a menu name, and a number displayed on the screen.
140 140 According to an exemplary embodiment, if only a part of the title of the object is displayed, and only a part of the converted text matches at least a part of the title of the displayed object, the controllerdetermines that the converted text corresponds to the title of the object. For example, if only “Stairway” is displayed of a content title “Stairway to Heaven,” and the converted text “stair” is input, the controllermay determine that the text corresponds to the title.
140 140 According to another exemplary embodiment, if only a part of one word included in the title of the object is displayed, and the converted text matches the whole one word, the controllerdetermines that the converted text corresponds to the title of the object. For example, if only “Stair” is displayed of the content title “Stairway to Heaven,” and the converted text “stairway” is input, the controllermay determine the converted text corresponds to the title.
140 140 130 140 130 140 140 If the controllerdetermines the converted text does not correspond to the title of the object, the controllerdetermines whether the converted text corresponds to a command stored in the storage. If the controllerdetermines the converted text corresponds to the command stored in the storage, the controllerperforms an operation corresponding to the command. Alternatively, the controllermay perform an operation that matches the command.
140 130 140 140 140 Also, if the controllerdetermines that the converted text does not correspond to the command stored in the storage, the controllerdetermines whether a meaning of the converted text is analyzable. If the controllerdetermines that the meaning of the converted text is analyzable, the controllermay analyze the meaning of the converted text and display a response message corresponding to the analysis result.
140 140 If the controllerdetermines the meaning of the converted text is not analyzable, the controllermay perform a search by using the converted text as a keyword.
140 140 140 100 As described above, the controllermay directly perform the work of analyzing the user voice signal and converting the user voice signal into converted text. However, according to other exemplary embodiments, the controllermay transmit the user voice signal to an external server apparatus, and the external server apparatus may convert the user voice signal into text. Also, the controllermay be provided with the converted text. The external server apparatus that converts a user voice signal into text may be referred to as a voice recognition apparatus for convenience of description. An operation of the display apparatusthat operates along with a voice recognition apparatus to convert a voice into a text will be described in detail in a subsequent exemplary embodiment.
130 100 130 130 140 The storageis an element that stores various types of modules for driving the display apparatus. The storagemay store a plurality of determination criteria and a plurality of commands for providing a voice recognition effect. For example, the storagemay store software including a voice conversion module, a text analysis module, a plurality of determination criteria, a control analysis criteria, a base module, a sensing module, a communication module, a presentation module, a web browser module, and a service module. The plurality of determination criteria may include whether converted text corresponds to a title of an object displayed on the screen, whether the converted text corresponds to a stored command, whether the converted text is grammatically analyzable, and whether the converted corresponds to a search keyword. The controllermay sequentially move through the plurality of determination criteria and apply the plurality of determination criteria sequentially in order to the converted text until a control operation is determined. In other words, the determination criteria represent a hierarchy of different tiers. That is, first it is determined whether the converted text corresponds to a displayed object in a first tier, and then if not, it is determined whether the converted text corresponds to a stored command in a second tier, and so on. After the control operation is determined, the control operation is performed.
130 100 100 100 130 100 According to an exemplary embodiment, the storagemay store at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus. A command may be stored in the storagethrough an input of the user. The command of the display apparatusis not limited thereto and may be various types of command.
100 1 FIG. The display apparatusindependently performs voice control inbut may operate along with the external server apparatus to perform the voice control.
2 FIG. 1 FIG. is a block diagram illustrating a detailed structure of the display apparatus ofaccording to an exemplary embodiment.
3 FIG. is a block diagram illustrating a software structure of a storage according to an exemplary embodiment.
2 FIG. 100 102 110 120 106 107 108 130 117 118 105 140 103 104 Referring to, the display apparatusincludes a display, the voice circuit, the voice converter, a communicator, an image receiver, an audio output circuit, the storage, an image processor, an audio processor, an input circuit, the controller, a speaker, and a remote control receiver.
2 FIG. 2 FIG. 100 In, the display apparatusis described as an apparatus having various functions such as a communication function, a broadcast receiving function, a video playback function, a display function, etc. to synthetically illustrate various types of elements. Therefore, according to exemplary embodiments, some of the elements ofmay be omitted or changed or other types of elements may be added.
102 107 117 The displaydisplays at least one of a video frame and various types of screens generated by a graphic processor (not shown). The video frame is formed by receiving image data from the image receiverand then processing the image data by the image processor.
106 106 The communicatoris an element that communicates with various types of external apparatuses or an external server. The communicatormay include various types of communication chips such as a WiFi chip, a Bluetooth chip, a near field communication (NFC) chip, a wireless communication chip, etc. Here, the WiFi chip, the Bluetooth chip, and the NFC chip respectively perform communications according to a WiFi method, a Bluetooth method, and an NFC method. The NFC chip refers to a chip that operates according to an NFC method using a band of 13.56 MHz among various radio frequency identification (RFID) frequency bands such as 135 kHz, 13.56 MHz, 433 MHz, 860˜960 MHz, 2.45 GHz, etc. If the WiFi chip or the Bluetooth chip is used, various types of connection information, such as a subsystem identification (SSID), a session key, etc., may be first transmitted and received to perform communication connections by using the various types of connection information and then transmit and receive various types of information. The wireless communication chip refers to a chip that perform communications according to various communication standards such as IEEE, Zigbee, 3rd Generation (3G), 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), etc.
106 106 According to an exemplary embodiment, the communicatormay transmit a user voice signal to a voice recognition apparatus and receive text, into which the user voice is converted, from the voice recognition apparatus. The communicatormay also store text information and search information desired by a user in an external server apparatus.
107 107 The image receiverreceives the image data through various sources. For example, the image receivermay receive broadcast data from an external broadcasting station or may receive image data from an external apparatus (for example, a digital versatile disc (DVD) apparatus).
108 118 108 The audio output circuitis an element that outputs various types of audio data processed by the audio processor, various types of notification sounds or voice messages. In particular, the audio output circuitmay output the user voice signal received from an external source.
130 100 130 130 131 132 133 134 1 134 2 130 100 3 FIG. 3 FIG. The storagestores various types of modules for driving the display apparatus. The modules stored in the storagewill now be described with reference to. As shown in, the storagemay store software including a voice conversion module, a text analysis module, a user interface (UI) framework, a plurality of determination criteria-, a plurality of commands-, a base module, a sensing module, a communication module, a presentation module, a web browser module, and a service module. The software modules stored in the storage, when executed by a microprocessor, cause the display apparatusto perform various actions.
130 100 100 100 134 2 130 100 According to an exemplary embodiment, the storagemay store at least one of a command for controlling power of the display apparatus, a command for controlling a channel of the display apparatus, and a command for controlling a volume of the display apparatus. The plurality of commands-may be stored in the storagethrough an input of the user. A command of the display apparatusis not limited to the above-described exemplary embodiment but may be various types of commands.
131 The voice conversion module, when executed by a processor, such as a microprocessor or microcontroller, converts a voice input by the user into text to output text information.
132 100 The text analysis module, when executed by a processor, such as a microprocessor or microcontroller, analyzes the converted text to perform an accurate function of the display apparatus.
100 Here, the base module, when executed by a processor, such as a microprocessor or microcontroller, processes signals transmitted from respective pieces of hardware included in the display apparatusand transmits the processed signals to an upper layer module. The sensing module, when executed by a processor, such as a microprocessor or microcontroller, collects information from various types of sensors, and analyzes and manages the collected information and may include a face recognition module, a voice recognition module, a motion recognition module, an NFC recognition module, etc. The presentation module, when executed by a processor, such as a microprocessor or microcontroller, configures a display screen and may include a multimedia module for playing back and outputting a multimedia content and a UI rendering module for processing a UI and a graphic. The communication module, when executed by a processor, such as a microprocessor or microcontroller, communicates with an external apparatus. The web browser module, when executed by a processor, such as a microprocessor or microcontroller, performs web browsing to access a web server. The service module includes various types of applications, which when executed by a processor, such as a microprocessor or microcontroller, provide various types of services.
2 FIG. 110 110 100 100 110 120 Returning to, the voice input circuitis an element that receives the voice of the user. In detail, the voice input circuitmay include a microphone and associated circuitry to directly receive the user's voice as sound and convert the sounds to an electric signal, or may include circuitry to receive an electric signal corresponding to the user's voice input to the display apparatusthrough a microphone that is connected to the display apparatusby wire or wireless. The voice input circuittransmits the signal corresponding to the user's voice to the voice converter.
120 The voice converterparses a waveform of a characteristic of the user's voice signal (i.e., a characteristic vector of the user's voice signal) to recognize words or a word string corresponding to the voice uttered by a user and outputs the recognized words as text information.
120 120 120 118 118 180 108 In detail, the voice convertermay recognize the words or the word string uttered by the user from the user voice signal by using at least one of various recognition algorithms such as a dynamic time warping method, a Hidden Markov model, a neural network, etc. and convert the recognized voice into text. For example, if the Hidden Markov model is used, the voice converterrespectively models a time change and a spectrum change of the user voice signal to detect a similar word from a stored language database (DB). Therefore, the voice convertermay output the detected word or words as text. The audio processoris an element that processes audio data. The audio processormay perform various types of processing, such as decoding, amplifying, noise filtering, etc., with respect to the audio data. The audio data processed by the audio processormay be output to the audio output circuit.
105 100 105 The input circuitreceives a user command for controlling an overall operation of the display apparatus. In particular, the input circuitmay receive a user command for executing a voice input mode, a user command for selecting a service that is to be performed, or the like.
105 105 100 The input circuitmay be realized as a touch panel, but this is only an example. Alternatively, for example, the input circuitmay be realized as another type of input device that may control the display apparatus, such as a remote controller, a pointing device, or the like.
140 100 130 140 The controllercontrols an overall operation of the display apparatusby using various types of programs and modules stored in the storage. The controllermay be one or more microprocessors.
140 110 140 140 110 120 140 130 140 100 The controllerperforms a control operation corresponding to the user voice input through the voice input circuit. The controllermay start a voice input mode according to a selection of the user. If the voice input mode starts, the controllermay activate the voice input circuitand the voice converterto receive the user's voice. If the user voice is input when a voice input mode is active, the controlleranalyzes an intention of the user by using a plurality of different determination criteria stored in the storage. The controllerdetermines a control operation according to the analysis result to perform an operation of the display apparatus.
140 140 140 4 6 FIGS.through In detail, the controllerdetermines whether the converted text corresponds to a title of an object in a displayed screen. If the converted text corresponds to the title of the object, the controllerperforms an operation corresponding to the object. For example, the controllermay perform an operation matching the object. In detail, the object may include at least one of a content title, an image title, a text icon, a menu name, and a number displayed on the screen. An operation of determining whether the text corresponds to the displayed object will now be described with reference to.
4 FIG. 100 100 140 10 100 120 140 1 410 12 1 410 is a view illustrating text that corresponds to a title displayed in a screen displayed on the display apparatus, or to a number displayed on a side of the display apparatus. The text may match or be equal to the title displayed or the number displayed. If the text into which the user voice signal is converted corresponds to the title displayed in a display screen, the controllerperforms a corresponding function. For example, if a userutters a voice “turn on”, the display apparatusconverts the input voice into text through the voice converter. The controllerdetermines that channelof “turn on” of a plurality of titles displayed on a screen of a TV corresponds to the text “turn on” into which the input voice is converted, to change channelof “BBB”, which is currently shown, to channelof “turn on”.
5 FIG. 100 140 10 100 120 510 530 510 140 510 According to another exemplary embodiment,illustrates text that corresponds to a plurality of icons displayed on a screen of the display apparatusand/or titles of the icons. The text may match or be equal to the plurality of icons displayed and/or the titles of the icons. If the text into which the user voice is converted corresponds to the plurality of icons displayed on the screen and the titles of the icons, the controllerperforms a corresponding function. For example, if the userutters a voice “more”, the display apparatusconverts the voice into text through the voice converter. Since a plurality of objectsthroughinclude an icon“more” corresponding to the text “more” into which the voice is converted, the controllerperforms a function of the icon“more”.
6 FIG. 100 140 10 1 3 100 100 120 140 1 610 610 690 According to another exemplary embodiment,illustrates text that corresponds to a plurality of menus displayed on a screen of the display apparatusand/or numbers displayed on upper sides of the menus. The text may match or be equal to the plurality of menus displayed and/or the numbers displayed. If the text into which the voice is converted corresponds to the plurality of menus displayed on the screen and/or the numbers displayed on the upper sides of the menus, the controllerperforms a corresponding function. For example, if the userutters a voice “numbermenu”when a list of a plurality menus is displayed on the display apparatus, and/or numbers are respectively displayed on upper sides of the menus, the display apparatusconverts the input voice into text through the voice converter. The controllerexecutes “numbermenu”corresponding to the converted text from a list of a plurality of menusthroughin which numbers are displayed.
7 FIG. 1 710 2 720 3 730 100 10 100 120 140 10 1 710 1 710 According to another exemplary embodiment, as shown in, if only a part of an object is displayed, and the text into which a voice is converted corresponds to at least a part of a title of the displayed object, a determination is made that the converted text corresponds to the title of the object. For example, if function execution objects,, andare displayed on the display apparatusas “THE LORD OF . . . ”, “AN APPLE IS . . . ”, and “MY NAME IS . . . ”, respectively, and the userutters a voice “Lord”, the display apparatusconverts the input voice into text through the voice converter. The controllerdetermines that the converted text “Lord” uttered by the usercorresponds to at least a part of functional execution objectthat is displayed, and performs a function of the function execution object.
8 FIG. 1 810 2 820 3 830 100 10 100 120 140 10 1 810 1 810 According to another exemplary embodiment, as shown in, if only a part of one of words of an object is displayed, and the converted text corresponds to the whole one word, a determination is made that the converted text corresponds to a title of the object. For example, if function execution objects of,, andare displayed on the display apparatusas “THE STO . . . ”, “THE APPL . . . ” and “THE HOU . . . ”, respectively, and the userutters a whole word “story”, the display apparatusconverts the input voice “story” into text through the voice converter. The controllerdetermines that the converted text “story” uttered by the usercorresponds to the title of the function execution object, and performs a function of the function execution object.
140 130 140 130 140 140 11 25 140 120 140 140 130 140 11 9 FIG. 9 FIG. According to an exemplary embodiment, if the text into which the user voice is converted does not correspond to the title of the object, the controllerdetermines whether the converted text corresponds to a command stored in the storage. If the controllerdetermines that the converted text corresponds to the command stored in the storage, the controllerperforms an operation corresponding to the command. Alternatively, the controllermay perform an operation matching the command. For example, referring to, if channelof a TV is shown as denoted by, and a voice “volume up” is input, the controllerconverts the input voice “volume up” into text through the voice converter. The controllerdetermines whether the converted text “volume up” corresponds to an object displayed on a screen of the TV. As shown in, since an object corresponding to the text “volume up” does not exist on the screen of the TV, the controllerdetermines that the converted text “volume up” corresponds to a command stored in the storage, and the controllerturns up a volume of the displayed channel.
130 140 140 140 10 11 25 140 120 140 25 100 140 130 130 140 130 140 140 145 11 100 10 FIG. Also, as described above, if an input voice does not correspond to a command stored in the storage, the controllerdetermines whether a meaning of the input voice is grammatically analyzable. If the controllerdetermines that the meaning of the input voice is grammatically analyzable, the controlleranalyzes the meaning of the input voice and displays a response message corresponding to the analysis result. For example, referring to, if the userinputs a voice “what time” when channelof the TV is shown on a screen, the controllerconverts the input voice “what time” into text through the voice converter. Here, the controllerdetermines whether the converted text corresponds to a title of an object displayed on the screenof the display apparatus. Since an object corresponding to the converted text “what time” is not displayed, the controllerdetermines whether a function corresponding to the converted text “what time” is stored in the storage. Assuming here that the converted text “what time” does not correspond to a function stored in the storage, the controllerdetermines whether the converted text “what time” is grammatically uttered, according to a criterion stored in the storage. If the controllerdetermines that the converted text “what time” is grammatically uttered, the controllermay display time information“It is 11:00 AM.” on a side of the screen on which channelis shown. In other words, the display apparatusmay display a response message corresponding to the analysis result on the screen.
140 140 10 140 100 130 140 140 140 140 155 156 157 140 155 156 157 10 110 155 140 155 157 10 140 10 11 FIG. 11 FIG.A 11 FIG.B 11 FIG.C 11 FIG.B 11 FIG.C According to another exemplary embodiment, as described above, if the controllerdetermines that the meaning of the input voice is not grammatically analyzable, the controllermay perform a search by using the converted text as a keyword. For example, referring to, if the userinputs a voice “AAA”, the controllerdetermines whether the converted text “AAA” corresponds to a title of an object displayed on a screen of the display apparatus. Assuming that the converted text “AAA” does not correspond to an object displayed on the screen or a command stored in the storage, the controllerdetermines if the converted text “AAA” is grammatically analyzable. Here, the controllerdetermines that the converted text “AAA” is not grammatically analyzable, and the controllermay perform a search by using the converted text “AAA” as a keyword. The controllermay perform a search in relation to the converted text “AAA” and display the search result on a screen, as shown in. According to an exemplary embodiment, if a plurality of search results,,, . . . are searched, the controllermay display the plurality of search results,,, . . . along with corresponding numbers. If the userinputs one of a plurality of lists through the voice circuit, i.e., inputs a voice “number one”, the controllermay display “AAA news broadcast time” along with a plurality of screen timesand. Here, an item selected by the usermay be displayed with a color, shape change, animation, etc. to be distinguished from other items. For example, in, the controllerhighlights “AAA news broadcast time”. In, if the userthen inputs a voice “select”, the selected items is executed. In response to the user inputting a voice “select” after “AAA news broadcast time” is highlighted in, the “AAA news broadcast time” of “8:00 AM”, “9:00 AM”, “2:00 AM”, and “6:00 AM” is shown, as in.
2 FIG. 140 109 111 113 112 101 109 111 113 112 101 111 112 130 109 111 Returning to, the controllerincludes a random access memory (RAM), a read only memory (ROM), a graphic processor, a main central processing unit (CPU), first through nth interfaces (not shown), and a bus. Here, the RAM, the ROM, the graphic processor (GPU), the main CPU, the first through nth interfaces, etc. may be connected to one another through the bus. The ROMstores a command set for booting a system, etc. If power is supplied through an input of a turn-on command, the main CPUcopies an operating system (O/S) stored in the storageinto the RAMand executes the O/S to boot the system according to the command stored in the ROM.
113 105 102 The GPUgenerates a screen including various types of objects, such as an icon, an image, a text, etc., by using an operator (not shown) and a renderer (not shown). The operator calculates attribute values, such as coordinate values, shapes, sizes, colors, etc. in which objects will be respectively displayed according to a layout of a screen, by using a control command received from the input circuit. The renderer generates a screen having various types of layouts including objects based on the attribute values calculated by the operator. The screen generated by the renderer is displayed in a display area of the display.
112 130 130 112 130 112 The main CPUaccesses the storageto perform booting by using the O/S stored in the storage. The main CPUalso performs various operations by using various types of programs, contents, data, etc. stored in the storage. The main CPUmay comprise at least one microprocessor and/or at least one microcontroller.
The first through nth interfaces are connected to various types of elements as described above. One of the first through nth interfaces may be a network interface that is connected to an external apparatus through a network.
140 106 130 In particular, the controllermay store text, into which a user voice signal is converted and which is received from a voice recognition apparatus through the communicator, in the storage.
105 140 140 110 120 100 If a voice recognition mode change command is input through the input circuit, the controllerexecutes a voice recognition mode. If the voice recognition mode is executed, the controllerconverts a voice of a user into text using the voice input circuitand the voice converteras described above to control the display apparatus.
103 118 The speakeroutputs audio data generated by the audio processor.
12 FIG. is a flowchart of a method of controlling a voice, according to an exemplary embodiment.
12 FIG. 100 1210 100 100 Referring to, if a voice input mode starts, the display apparatusreceives a user voice from a user in operation S. As described above, the user voice may be input through a microphone installed in a main body of the display apparatus, or through a remote controller or microphones installed in other external apparatus and then may be transmitted to the display apparatus.
1220 120 100 100 The input user voice is converted into text in operation S. The conversion of the user voice into text may be performed by the voice converterof the display apparatus, or by an external server apparatus separately installed outside the display apparatus.
1230 140 100 134 1 130 In operation S, the controllerof the display apparatussequentially changes and applies the plurality of determination criteria-stored in the storageto the converted text.
100 140 1240 If the converted text corresponds to an object displayed on a screen of the display apparatus, the controllerdetermines a control operation corresponding to the object in operation S.
As described above, the object displayed on the screen may be at least one of a content title, an image title, a text icon, a menu name, and a number.
The operation is then performed, and the process ends.
13 FIG. is a flowchart of a method of controlling a voice, according to another exemplary embodiment.
13 FIG. 100 1310 100 100 Referring to, if a voice input mode starts, the display apparatusreceives a voice of a user in operation S. As described above, the user voice may be input through a microphone installed in a main body of the display apparatus, or through a remote controller or microphones installed in other external apparatuses and then transmitted to the display apparatus.
1320 120 100 100 100 106 The input user voice is converted into text in operation S. The conversion of the user voice into text may be performed by the voice converterof the display apparatus. However, according to another exemplary embodiment, the display apparatusmay transmit the user voice to an external server apparatus, the external server apparatus may convert the user voice into text, and the display apparatusmay receive the converted text through the communicator.
1330 100 100 In operation S, the display apparatusdetermines whether the converted text corresponds to a title of an object displayed on a screen of the display apparatus. The object displayed on the screen may include at least one of a content title, an image title, a text icon, a menu name, and a number and may be variously realized according to types of the object.
1330 100 1335 1 100 100 12 1 4 FIG. If it is determined that the text corresponds to the title of the object displayed on the screen (operation S, Y), the display apparatusperforms an operation corresponding to the object as a control operation in operation S. For example, if the user inputs a voice “turn on”, and the converted text “turn on” corresponds to “turn on channel” of a plurality of titles displayed on the screen of the display apparatus, the display apparatuschanges channelof “BBB” that is currently shown into channel. (See).
1330 100 130 1340 If it is determined that the text does not correspond to the title of the object displayed on the screen (operation S, N), the display apparatusdetermines whether the text corresponds to a command stored in the storagein operation S.
130 Here, the storagemay designate and store voice commands respectively corresponding to various operations such as turn-on, turn-off, volume-up, volume-down, etc.
130 1340 100 1345 130 140 9 FIG. If it is determined that the text corresponds to the command stored in the storage(operation S, Y), the display apparatusdetermines an operation corresponding to the command as a control operation in operation S. For example, if the user utters a voice “volume up”, and a command corresponding to the voice “volume up” is stored in the storage, the controllerturns up a volume of a corresponding channel. (See).
130 1340 100 1350 If it is determined that the text does not correspond to the command stored in the storage(operation S, N), the display apparatusdetermines whether a meaning of the text is analyzable in operation S.
1350 100 1355 100 100 10 FIG. If it is determined that the meaning of the text is analyzable (operation S, Y), the display apparatusdisplays a response message corresponding to the analysis result in operation S. For example, if the user utters a voice “How's the weather today?”, the display apparatusmay grammatically analyze the converted text “How's the weather today?” and display information about the weather on the screen of the display apparatus. (See also).
1350 100 1360 100 11 11 FIGS.A-C If it is determined that the meaning of the text is not analyzable (operation S, N), the display apparatusperforms a search by using the converted text as a keyword in operation S. For example, as in the above-described exemplary embodiment, if the user utters “AAA”, the display apparatusmay perform a search using the converted text “AAA” as a keyword and display the search result in a result display area. Here, if a plurality of search results for the keyword “AAA” exist, one of the plurality of search results selected by the user may be realized and displayed according to various methods such as a color, a shape, an animation, etc. (See).
14 FIG. is a view illustrating a structure of a voice control system according to an exemplary embodiment.
14 FIG. 600 350 300 100 In detail, referring to, a voice control systemincludes a voice recognition apparatus, a server apparatus, and the display apparatus.
100 350 300 140 The display apparatusmay include a client module that may operate along with the voice recognition apparatusand the server apparatus. If a voice input mode starts, the controllermay execute the client module to perform a control operation corresponding to a voice input.
140 100 350 106 350 100 100 2 FIG. In detail, if a user voice is input, the controllerof the display apparatusmay transmit the user voice to the voice recognition apparatusthrough the communicator. (See). The voice recognition apparatusdenotes a server apparatus that converts the user voice transmitted through the display apparatusinto text and returns the converted text to the display apparatus.
350 100 100 134 1 130 If the converted text is returned from the voice recognition apparatusto the display apparatus, the display apparatussequentially applies the plurality of determination criteria-stored in the storageto perform an operation corresponding to the converted text.
100 300 300 300 100 According to an exemplary embodiment, the display apparatusmay provide the server apparatuswith the converted text into which the user voice is converted. The server apparatusmay search a database DB thereof or other server apparatuses for information corresponding to the provided converted text. The server apparatusmay feed the search result back to the display apparatus.
According to various exemplary embodiments, even if a user inputs a simple keyword having no verb, an accurate operation corresponding to an intention of the user may be processed.
The foregoing exemplary embodiments and advantages are merely exemplary and are not to be construed as limiting. The present teaching can be readily applied to other types of apparatuses. Also, the description of the exemplary embodiments is intended to be illustrative, and not to limit the scope of the claims, and many alternatives, modifications, and variations will be apparent to those skilled in the art.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
September 28, 2023
June 30, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.