A device includes memory configured to store audio data and one or more processors configured to obtain the audio data captured by a microphone of a wearable device. The instructions further cause the one or more processors to determine, based on one or more signals exchanged between the wearable device and a reference device, directionality information indicative of a direction of the microphone relative to the reference device. The instructions also cause the one or more processors to process the audio data based on the directionality information to generate spatial audio data.
Legal claims defining the scope of protection, as filed with the USPTO.
memory configured to store audio data; and obtain the audio data captured by a microphone of a wearable device; determine, based on one or more wireless communication signals exchanged between the wearable device and a reference device, directionality information indicative of a direction of the microphone relative to the reference device; process the audio data based on the directionality information to generate spatial audio data that corresponds to the audio data coming from the direction of the microphone relative to the reference device; and transmit the audio data to an output device distinct from the wearable device and the reference device. one or more processors configured to: . A device comprising:
claim 1 . The device of, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and wherein the one or more processors are configured to update the spatial audio data over time to represent movement of the wearable device relative to the reference device as movement of the source of the sound.
claim 1 . The device of, wherein the one or more wireless communication signals include encoded data, wherein the one or more processors are configured to receive the one or more wireless communication signals and decode the encoded data to generate the audio data.
claim 1 . The device of, wherein the one or more processors are configured to obtain the audio data from one or more data packets of the one or more wireless communication signals.
claim 1 . The device of, wherein the one or more processors are configured to determine the directionality information based on an angle of arrival of the one or more wireless communication signals.
claim 1 . The device of, further comprising one or more antennas configured to transmit a signal of the one or more wireless communication signals, to receive a signal of the one or more wireless communication signals, or both.
claim 1 . The device of, wherein the one or more processors are further configured to determine, based on a received signal strength of the one or more wireless communication signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
claim 1 . The device of, wherein the spatial audio data includes ambisonics data.
claim 1 . The device of, further comprising a camera coupled to the one or more processors and configured to capture video data, wherein the one or more processors are configured to process the video data in conjunction with the spatial audio data and to encode the video data and the spatial audio data for communication to another device.
claim 1 . The device of, further comprising a second microphone coupled to the one or more processors, wherein the one or more processors are configured to modify the audio data based on sound captured at the second microphone.
claim 10 . The device of, wherein the one or more processors are configured to modify the audio data to de-emphasize, in the spatial audio data, audio components that are present in both the audio data and in the sound captured at the second microphone.
claim 1 . The device of, wherein the one or more processors and the memory are integrated within the reference device.
claim 1 . The device of, wherein the one or more processors and the memory are integrated within the wearable device.
claim 1 . The device of, wherein the one or more processors and the memory are integrated into at least one of a smart speaker, a speaker bar, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a communication device, an internet-of-things (IoT) device, an extended reality (XR) device, a base station, or a mobile device.
claim 1 . The device of, wherein the wearable device corresponds to or includes a headset device or one or more earbuds.
obtaining, at one or more processors, audio data captured by a microphone of a wearable device; determining, by the one or more processors, directionality information indicative of a direction between the microphone and a reference device based on one or more wireless communication signals exchanged between the wearable device and the reference device; and generating, at the one or more processors, spatial audio data based on the audio data and the directionality information that corresponds to the audio data coming from the direction of the microphone relative to the reference device; and transmitting the audio data to an output device distinct from the wearable device and the reference device. . A method comprising:
claim 16 after determining the directionality information, determining updated directionality information; and generating updated spatial audio data, wherein the updated spatial audio data represents movement over time of the wearable device relative to the reference device as movement of the source of the sound. . The method of, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and further comprising:
claim 16 . The method of, wherein the one or more wireless communication signals include encoded data, and further comprising receiving the one or more wireless communication signals and decoding the encoded data to generate the audio data.
claim 16 . The method of, wherein the directionality information is based on an angle of arrival of the one or more wireless communication signals.
claim 16 . The method of, further comprising determining, based on a received signal strength of the one or more wireless communication signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
claim 16 . The method of, wherein the spatial audio data includes ambisonics data.
claim 16 obtaining video data associated with the audio data; processing the video data in conjunction with the spatial audio data; and encoding the video data and the spatial audio data for transmission or storage. . The method of, further comprising:
claim 16 . The method of, further comprising modifying the audio data based on sound captured at a second microphone.
obtain audio data captured by a microphone of a wearable device; determine directionality information indicative of a direction between the microphone and a reference device based on one or more wireless communication signals exchanged between the wearable device and the reference device; generate spatial audio data based on the audio data and the directionality information that corresponds to the audio data coming from the direction of the microphone relative to the reference device; and transmit the audio data to an output device distinct from the wearable device and the reference device. . A non-transitory computer-readable device storing instructions that are executable by one or more processors to cause the one or more processors to:
claim 24 after determining the directionality information, determine updated directionality information; and generate updated spatial audio data based on the updated directionality information, wherein the updated spatial audio data represents movement over time of the wearable device relative to the reference device as movement of the source of the sound. . The non-transitory computer-readable device of, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and wherein the instructions are further executable to:
claim 24 . The non-transitory computer-readable device of, wherein the one or more wireless communication signals include encoded data, wherein the instructions are further executable to decode the encoded data to generate the audio data.
claim 24 . The non-transitory computer-readable device of, wherein the directionality information is based on an angle of arrival of the one or more wireless communication signals.
claim 24 . The non-transitory computer-readable device of, wherein the instructions are further executable to determine, based on a received signal strength of the one or more wireless communication signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
claim 24 . The non-transitory computer-readable device of, wherein the spatial audio data includes ambisonics data.
means for obtaining audio data captured by a microphone of a wearable device; means for determining directionality information indicative of a direction between the microphone and a reference device based on one or more wireless communication signals exchanged between the wearable device and the reference device; means for generating spatial audio data based on the audio data and the directionality information that corresponds to the audio data coming from the direction of the microphone relative to the reference device; and means for transmitting the audio data to an output device distinct from the wearable device and the reference device. . An apparatus comprising:
Complete technical specification and implementation details from the patent document.
The present disclosure is generally related to audio processing, and in particular to generation of spatial audio data.
Advances in technology have resulted in smaller and more powerful computing devices. For example, there currently exist a variety of portable personal computing devices, including wireless telephones such as mobile and smart phones, tablets and laptop computers that are small, lightweight, and easily carried by users. These devices can communicate voice and data packets over wireless networks. Further, many such devices incorporate additional functionality such as a digital still camera, a digital video camera, a digital recorder, and an audio file player. Also, such devices can process executable instructions, including software applications, such as a web browser application, that can be used to access the Internet. As such, these devices can include significant computing capabilities.
Such computing devices often incorporate functionality associated with communications and media generation. For example, such computing devices often support features such as voice calls, video calls, multi-party calls (e.g., teleconferencing and videoconferencing), gaming, multimedia capture, or combinations thereof.
During use of many of these features, audio is captured locally and saved or sent to another device, and a richer audio experience can be provided using spatial audio. Commonly, audio data is generated, stored, and/or rendered as channel-based audio (such as mono-channel audio, stereo-channel audio, 5.1 channel audio, etc.), in which each channel represents an output stream. In contrast, spatial audio data includes data representing objects or sound sources and output channels are generated as part of rendering the spatial audio data. For example, spatial audio can be rendered so that speech represented in the output audio sounds like it is coming from the direction of a person who is speaking. Unfortunately, capturing spatial audio data generally requires the use of complicated, special-purpose microphone arrays. Because of the complexity and expense of such microphone arrays, it is challenging and cost prohibitive to generate spatial audio for everyday use.
In a particular aspect, a device includes memory configured to store audio data and one or more processors configured to obtain the audio data captured by a microphone of a wearable device. The instructions further cause the one or more processors to determine, based on one or more signals exchanged between the wearable device and a reference device, directionality information indicative of a direction of the microphone relative to the reference device. The instructions also cause the one or more processors to process the audio data based on the directionality information to generate spatial audio data.
In a particular aspect, a method includes obtaining, at one or more processors, audio data captured by a microphone of a wearable device. The method also includes determining, by the one or more processors, directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device. The method further includes generating, at the one or more processors, spatial audio data based on the audio data and the directionality information.
In a particular aspect, a non-transitory computer-readable device stores instructions that are executable by one or more processors to cause the one or more processors to obtain audio data captured by a microphone of a wearable device. The instructions further cause the one or more processors to determine directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device. The instructions also cause the one or more processors to generate spatial audio data based on the audio data and the directionality information.
In a particular aspect, an apparatus includes means for obtaining audio data captured by a microphone of a wearable device. The apparatus also includes means for determining directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device. The apparatus further includes means for generating spatial audio data based on the audio data and the directionality information.
Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims.
8 FIG. Complicated, special-purpose microphone arrays are usually used to capture spatial audio. For example, the typical set up to capture sound to generate a first-order ambisonics representation of the sound includes at least four microphones including one omnidirectional microphone and threemicrophones arranged to have lobes along mutually orthogonal axes. Because of the complexity and expense of such microphone arrays, it is challenging and cost prohibitive to generate spatial audio for everyday use, such as for use during teleconferences.
Although directional microphone arrays are complex and not in common use, wearable devices that include a microphone (or perhaps a few microphones) arranged to capture speech of a user are quite common and inexpensive. For example, wireless headphones, earbuds, headsets, extended reality (XR) devices (e.g., XR headsets and XR glasses), etc. are increasingly common and often used for audio capture, among other things. Aspects disclosed herein enable use of such devices (e.g., wireless wearable devices that include one or more microphones) to generate spatial audio data without the use of special-purpose directional microphone arrays.
For example, a wearable device with a microphone can be used in conjunction with a reference device. The wearable device can be configured such that the microphone is disposed at a fixed position relative to the user's mouth. To illustrate, the wearable device can be head mounted, such that when the user's head moves, the microphone also moves and remains stationary relative to the user's mouth. Alternatively, the wearable device can be configured such that movement of the user's mouth relative to the microphone is constrained. To illustrate, the wearable device can be chest mounted or coupled to a limb of the user, such that when the user's head moves, the microphone moves a small amount relative to the user's mouth.
The reference device and the wearable device can exchange signals, such as advertisement packets, beacon signals, or signals encoding audio data captured by the wearable device, and the reference device can use the signals to determine directionality information indicating a direction from the reference device to the wearable device. For example, the reference device can determine the angle of arrival of signals from the wearable device. As another example, the reference device can determine range information indicating a distance between the wearable device and the reference device.
The directionality information and the audio data captured by the wearable device can be used to generate spatial audio data, such as first-order ambisonics data. For example, the spatial audio data may represent the audio data as coming from a sound source that is at the location of the wearable device relative to the reference device. As the user moves relative to the reference device, a sound source represented in the spatial audio data moves due to the changing position of the wearable device relative to the reference device.
In some embodiments, the audio data can be sent from the wearable device to the reference device (or to another device) via the same signals as are used to determine the position of the wearable device relative to the reference device. In such examples, no additional processing burden is placed on the wearable device to enable generation of spatial audio data thereby conserving resources (e.g., battery power) of the wearable device and enabling the use of commonly available wearable devices to generate spatial audio data rather than special-purpose microphone arrays.
One problem with widespread use of 3D audio capture is the complexity and expense of 3D audio capture equipment. For example, many users capture audio using portable computing devices, such as smartphones, which are often too small to house a 3D audio capture microphone array. Further, many users use wireless headphones, earbuds, or other similar devices to capture audio data. Such devices are generally configured specifically for capturing speech of the user; accordingly, the microphones of such devices are usually configured to remain at a fixed position relative to the user's mouth to improve the quality of captured speech audio. One benefit of such devices is that they allow the user to move about the environment while capturing high-quality audio. A user's movement about the environment represents a circumstance in which 3D audio capture would be useful (e.g., to adapt audio output to represent the user's movement); however, it is problematic to capture 3D audio in this situation because of the fixed microphone position of a wireless device worn by the user, available microphones of a smartphone device, and the cost and complexity of special-purpose 3D audio capture equipment.
Disclosed embodiments provide a solution to these and other problems by generating spatial audio data using audio captured at a wearable device and directionality information based on signals exchanged between the wearable device and a reference device. For example, the user can move around the environment while speaking, while a reference device (such as a smartphone) exchanges signals with the wearable device. The distance, direction, or both, between the reference device and the wearable device are determined based on the signal exchange, and the audio data captured at the wearable device is modified based on the distance, direction, or both, to generate spatial audio data. For example, direction information can be determined based on the angle of arrival of the exchanged signals, distance information can be determined based on signal strength indicators associated with the exchanged signals, or both. Thus, one technical advantage of the disclosed embodiments is that low-cost and readily available audio capture equipment can be used to capture spatial audio data (e.g., 3D audio data). For example, the audio data can be processed to generate Ambisonics audio data (such as a 1st or 2nd order ambisonics representation of the audio captured at the wearable device).
1 FIG. 1 FIG. 192 190 192 190 192 190 Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terminology is used for the purpose of describing particular embodiments only and is not intended to be limiting of embodiments. For example, the singular forms “a.” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, some features described herein are singular in some embodiments and plural in other embodiments. To illustrate,depicts a deviceincluding one or more processors (“processor(s)”of), which indicates that in some embodiments the deviceincludes a single processorand in other embodiments the deviceincludes multiple processors. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural (as indicated by “(s)”) unless aspects related to multiple of the features are being described.
1 FIG. 130 130 130 130 130 In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and/or logically distinct, the same reference number is used for each, and the different instances are distinguished by addition of a letter to the reference number. When the features as a group or a type are referred to herein, e.g., when no particular one of the features is being referenced, the reference number is used without a distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, referring to, multiple locations are illustrated and associated with reference numbersA,B, andC. When referring to a particular one of these locations, such as a locationA, the distinguishing letter “A” is used. However, when referring to any arbitrary one of these locations or to these locations as a group, the reference numberis used without a distinguishing letter.
As used herein, the terms “comprise.” “comprises,” and “comprising” may be used interchangeably with “include.” “includes,” or “including.” Additionally, the term “wherein” may be used interchangeably with “where.” As used herein, “exemplary” indicates an example, an embodiment, and/or an aspect, and should not be construed as limiting or as indicating a preference or a preferred embodiment. As used herein, an ordinal term (e.g., “first.” “second.” “third.” etc.) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element with respect to another element, but rather merely distinguishes the element from another element having a same name (but for use of the ordinal term). As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.
As used herein, “coupled” may include “communicatively coupled.” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combinations thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof), etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some embodiments, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.
In the present disclosure, terms such as “determining,” “calculating,” “estimating,” “shifting,” “adjusting,” etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “generating,” “calculating,” “estimating,” “using,” “selecting,” “accessing,” and “determining” may be used interchangeably. For example, “generating,” “calculating,” “estimating,” or “determining” a parameter (or a signal) may refer to actively generating, estimating, calculating, or determining the parameter (or the signal) or may refer to using, selecting, or accessing the parameter (or signal) that is already generated, such as by another component or device.
1 FIG. 1 FIG. 100 180 100 110 120 150 100 180 162 110 110 112 180 150 112 112 is a diagram of a particular illustrative aspect of a systemoperable to generate spatial audio datain accordance with some examples of the present disclosure. The systemincludes a wearable devicethat includes one or more microphones(“mic(s)” in) and a reference device. The systemis configured to generate the spatial audio databased on audio datacaptured at the wearable devicewhile the wearable deviceis worn by a user. In a particular aspect, the spatial audio datarepresents relative positions of the reference deviceand the userand such relative positions may change over time (e.g., due to movement of the user).
110 112 182 162 110 120 182 112 110 120 182 112 112 150 112 150 120 182 112 110 120 112 110 112 120 112 150 The wearable deviceis configured to be worn by the userto capture soundto generate the audio data. For example, the wearable devicecan include a head-mounted device (e.g., a headset, earbuds, glasses, etc.) that includes the microphone(s)configured to capture the soundassociated with speech of the user. In other examples, the wearable deviceis configured to be worn on another portion of the user's body, such as on the user's chest or arm. In some embodiments, at least one of the microphone(s)is operable to capture soundgenerated by the userwhile the usermoves relative to the reference device. As compared to movement of the userrelative to the reference device, the microphone(s)remain stationary (or nearly stationary) relative to a source of soundgenerated by the user. For example, when the wearable deviceis a head-mounted device, the microphone(s)can be disposed at a fixed location (during use) relative to the user's mouth to capture speech of the user. As another example, when the wearable deviceis worn on the chest or a limb of the user, the microphone(s)can move during audio capture (e.g., due to turning of the user's head); however, such motion results in a relatively small position change as compared to movement of the userrelative to the reference device, and as such, can be ignored in some embodiments.
150 184 110 160 184 150 152 154 152 184 154 160 184 1 FIG. The reference deviceis configured to exchange one or more signalswith the wearable deviceand to generate directionality informationbased on the signal(s). For example, in, the reference deviceincludes a communication systemcoupled to one or more antennas. The communication systemis configured to send and/or receive the signalsvia the antenna(s)and to determine the directionality informationbased on the signals.
150 150 In some embodiments, the reference devicemay be integrated within a wireless access point configured to support a wireless local area network (WLAN), such as a WIFI® network (WI-FI® is a registered trademark of the Wi-Fi Alliance Corp., a California corporation). In other embodiments, the reference deviceis integrated within a portable communication device, such as a smart phone. To illustrate, the portable communication device may be configured to support a WLAN or a personal area network (PAN), such as one or more Bluetooth communication links (BLUETOOTH® is a registered trademark of Bluetooth SIG, Inc., a Delaware Corporation).
184 154 150 150 110 150 110 150 110 150 150 160 In some embodiments, ranging information, angle of arrival information, or both, can be determined based on a received signal strength of one or more signalsat particular antennas. Additionally, or alternatively, the reference devicemay be configured to determine a direction (e.g., an angle of arrival) using beamforming techniques, enabling the reference deviceto use multilateration techniques to determine location coordinates of the wearable device. Further, in some embodiments, the reference device, the wearable device, or both, are within a coverage area of several access points of a WLAN (e.g., part of a mesh network), and one or more access points may determine location information associated with the reference device, the wearable device, or both, and exchange the location information with the reference deviceto enable the reference deviceto determine more accurate and/or more precise directionality information.
150 184 160 150 160 184 150 184 150 160 In some implementations, the reference devicecan use phase-based ranging (based on the signals) to estimate the directionality information. One example of a phase-based ranging technique is high-accuracy distance measurement (HADM) based on Bluetooth Low-Energy (BLE) transmissions. Additionally, or alternatively, the reference devicecan use signal strength-based techniques to estimate the directionality information. To illustrate, the signalscan include a transmission power indicator, and the reference devicecan compare the transmission power indicator to a signal strength of the signalas received at the reference deviceto estimate ranging information. Multilateration based on the received signal strength or received signal strength fingerprinting for spatially diverse antennas can be used to estimate the directionality information.
110 184 150 160 184 184 150 154 184 184 154 150 184 184 184 184 110 150 184 154 In some embodiments, the wearable devicesends the signalsperiodically or occasionally, and the reference devicedetermines the directionality informationbased on an angle of arrival of the signals, ranging information associated with the signals, or both. For example, the reference devicecan include two or more antennas, and the angle of arrival of the signalscan be determined based on phases of waveforms of the signalsas received at the two or more antennas. In some such embodiments, the reference devicecan determine a received signal strength of the signalsand compare the received signal strength to a transmitted signal strength of the signalsto determine ranging information. In such embodiments, the transmitted signal strength of the signalscan be indicated in a data field of the signalsor can be determined based on prior agreement (e.g., based on a communication protocol specification or based on communication link set up data exchange between the wearable deviceand the reference device). In some embodiments, the ranging information can be determined based on multiple factors, such as differences in angle of arrival of the signalsat multiple spatially diverse antennasand signal strength information.
150 184 110 160 184 184 110 160 150 In some embodiments, the reference devicesends the signalsperiodically or occasionally, and the wearable devicedetermines the directionality informationbased on an angle of arrival of the signals, ranging information associated with the signals, or both, using similar techniques to those described above. In such embodiments, the wearable devicecan send the directionality informationto the reference device.
110 162 184 150 160 184 162 192 150 184 162 110 162 184 150 184 184 150 192 162 In some embodiments, the wearable deviceencodes the audio datato generate a stream of data packets, and transmits the data packets via the signals. In such embodiments, the reference devicecan determine the directionality informationbased on signals (e.g., the signals) that encode the audio data. A device, the reference device, or both, may be configured to receive encoded data representing the audio data via the signalsand to decode the encoded data to generate the audio data. For example, the wearable devicecan generate data packets that include data representing the audio dataand can transmit the data packets via the signalsusing a wireless protocol, such as a BLUETOOTH® protocol, a WIFI® protocol, or another local area or personal area wireless communication protocol. In this example, the reference devicecan receive the signalsand determine the directionality information based on the signals. The reference device, the device, or both, can also process the data packets to reconstruct the audio data.
1 FIG. 1 FIG. 100 192 190 180 162 160 190 140 150 130 110 140 180 In, the systemincludes the device, which includes one or more processorsconfigured to generate the spatial audio databased on the audio dataand the directionality information. For example, the processor(s)ofinclude a spatial audio generatorthat generates the spatial audio data as though the reference devicewere a spatial audio recording device (e.g., a multimicrophone array) that was capturing sound from a sound source at a locationof the wearable device. To illustrate, the spatial audio generatorcan include an ambisonics encoder that is configured to generate ambisonics data corresponding to the spatial audio data.
100 112 150 112 130 112 120 110 182 110 182 162 182 110 184 162 182 130 During operation of the systemin one example, the usercan walk about in a room or other space in which the reference deviceis located while speaking. In this example, at a first time (TO), the useris at a locationA. As the userspeaks, the microphone(s)of the wearable devicecapture soundA representing the user's speech. The wearable deviceprocesses the soundA to generate a portion of the audio datathat represents the soundA. The wearable devicetransmits signalsA representing the portion of the audio datarepresenting the soundA from the locationA.
150 184 184 130 160 130 150 184 130 160 The reference devicereceives the signalsA and determines an angle of arrival of the signalsA to determine a direction to the locationA. Directionality informationfor time TO includes an indication of the direction to the locationA. Optionally, in some embodiments, the reference devicealso determines range information based on a signal strength of the signalsA (e.g., based on a received signal strength and a transmitted signal strength indicator). The range information indicates a distance to the locationA. In such embodiments, directionality informationfor time TO also includes the range information.
140 162 182 160 180 140 162 182 150 130 140 162 182 110 150 192 184 140 162 182 150 The spatial audio generatorreceives the portion of the audio datarepresenting the soundA and the directionality informationfor time TO, and determines a portion of the spatial audio datafor the time TO. For example, the spatial audio generatorcan treat the portion of the audio datarepresenting the soundA as though it were captured by an ambisonics microphone array positioned at the reference devicefrom a sound source at the locationA. In some embodiments, the spatial audio generatorreceives the portion of the audio datarepresenting the soundA from the wearable device(e.g., both the reference deviceand the devicereceive the signalsA). In some embodiments, the spatial audio generatorreceives the portion of the audio datarepresenting the soundA from the reference device.
1 112 130 112 130 120 110 182 110 162 182 184 130 At a time T(after time TO), the userhas moved to a locationB. As the userspeaks at the locationB, the microphone(s)of the wearable devicecapture soundB representing the user's speech. The wearable devicesends a portion of the audio datarepresenting the soundB via signalsB from the locationB.
150 184 160 1 184 184 140 162 182 160 1 180 1 140 162 182 150 130 The reference devicereceives the signalsB and determines the directionality informationfor time Tbased on the angle of arrival of the signalsB, signal strength information associated with the signalsB, or both. The spatial audio generatorreceives the portion of the audio datarepresenting the soundB and the directionality informationfor time T, and determines a portion of the spatial audio datafor the time T. For example, the spatial audio generatorcan treat the portion of the audio datarepresenting the soundB as though it were captured by an ambisonics microphone array positioned at the reference devicefrom a sound source at the locationB.
2 1 112 130 112 130 120 110 182 110 162 182 184 130 Likewise, at a time T(after time T), the userhas moved to a locationC. As the userspeaks at the locationC, the microphone(s)of the wearable devicecapture soundC representing the user's speech. The wearable devicesends a portion of the audio datarepresenting the soundC via signalsC from the locationC.
150 184 160 2 184 184 140 162 182 160 2 180 2 140 162 182 150 130 The reference devicereceives the signalsC and determines the directionality informationfor time Tbased on the angle of arrival of the signalsC, signal strength information associated with the signalsC, or both. The spatial audio generatorreceives the portion of the audio datarepresenting the soundC and the directionality informationfor time T, and determines a portion of the spatial audio datafor the time T. For example, the spatial audio generatorcan treat the portion of the audio datarepresenting the soundC as though it were captured by an ambisonics microphone array positioned at the reference devicefrom a sound source at the locationC.
100 180 100 112 130 110 110 162 100 180 8 FIG. As the example above illustrates, the systemis able to generate spatial audio data(such as ambisonics data) using a less complex microphone arrangement than is conventional for spatial audio capture. To illustrate, first-order ambisonics audio capture typically uses a multimicrophone array that includes special purpose microphones. As one example, an ambisonics microphone array can include an omnidirectional microphone and threemicrophones, which are carefully positioned to have orthogonal main lobes. In contrast, the systemcan capture first-order ambisonics data representing speech of the userfrom various locationsusing a single microphone on the wearable device. In some embodiments, the wearable devicecan include additional microphones, which can be used to filter or otherwise modify the audio datato emphasize or enhance audio components associated with target sounds (e.g., speech), to de-emphasize audio components associated with non-target sounds (e.g., noise), or both. Thus, an advantage of the systemis generation of spatial audio datawithout special-purpose audio capture equipment.
1 FIG. 100 110 150 192 Although particular communication protocols and corresponding networks are described with reference to, in other embodiments, the systemcan use other types of communication protocols and/or networks. To illustrate, data exchange between the wearable device, the reference device, the device, and/or other devices (not shown) can use wide area wireless connection, often referred to as cellular connection or a mobile data connection, such as a connection conforming to a cellular voice and data network protocol from a 3rd Generation Partnership Project (3GPP) standards organization (e.g., a 3G, 4G, or 5G connection).
1 FIG. 110 112 110 112 110 112 112 182 Althoughillustrates the wearable deviceas worn by a person (e.g., the user), in other embodiments, the wearable devicecan be mounted on, coupled to, or integrated within another device that is associated with the user. For example, the wearable devicecan be mounted on, coupled to, or integrated within a vehicle (e.g., an unmanned aerial vehicle) that uses a “follow-me” feature to track and follow the useror in which the useris riding. As another example, the wearable device can be mounted on, coupled to, or integrated within a non-human user, such as an animal, a robot, a vehicle, etc. and used to capture soundsnear the non-human user.
150 192 110 In various embodiments, the reference device, the device, or both, include, correspond to, or are included within a smart speaker, a speaker bar, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a communication device, an internet-of-things (IoT) device, an extended reality (XR) device, a base station, or a mobile device. Additionally, or alternatively, in various embodiments, the wearable deviceincludes, corresponds to, or is included within a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, extended reality glasses, one or more earbuds, a wireless microphone device (e.g., a lapel microphone), etc.
2 FIG. 1 FIG. 2 FIG. 1 FIG. 2 FIG. 2 FIG. 200 100 150 192 150 190 140 150 is a block diagram of a systemthat includes or corresponds to particular illustrative aspects of the systemofin accordance with some examples of the present disclosure. In the example illustrated in, the reference deviceand the deviceofare combined. For example, the reference deviceofincludes the processor(s), which include the spatial audio generator.also illustrates additional details and optional components of the reference device.
2 FIG. 2 FIG. 2 FIG. 110 120 182 110 110 252 254 252 184 254 110 284 162 182 252 284 184 In, the wearable deviceincludes the microphone(s)to capture sound, such as speech of a user wearing the wearable device. The wearable deviceofalso includes a communication systemcoupled to one or more antennas. The communication systemis configured to send the signal(s)via the antenna(s)according to a wireless communication protocol, such as a BLUETOOTH® communication protocol, a WIFI® communication protocol, etc. In the example of, the wearable deviceis configured to generate data packetsthat include audio datarepresenting the sound, and the communication systemis configured to send the data packetsvia the signals.
2 FIG. 2 FIG. 150 154 152 150 190 140 150 210 190 210 162 160 180 150 180 212 190 In, the reference deviceincludes the antenna(s)coupled to the communication system. Further, as noted above, the reference deviceincludes the processor(s), which include the spatial audio generator. In, the reference devicealso includes a memorycoupled to the processor(s). The memoryis configured to store, for example, the audio data, the directionality information, the spatial audio data, other data used by the reference deviceto generate the spatial audio data, instructionsexecutable by the processor(s), or any combination thereof.
1 FIG. 2 FIG. 154 184 152 160 184 152 242 242 150 110 184 184 154 184 184 150 110 As described with reference to, the antenna(s)are configured to receive the signal(s), and the communication systemis configured to determine directionality informationbased on the signal(s). For example, the communication systemofincludes a direction detector. The direction detectoris configured to determine a direction from the reference deviceto the wearable devicebased on the signal(s). To illustrate, differences in the phases of the signal(s)when received at two or more of the antenna(s)can be used to determine an angle of arrival of the signal(s). In this example, the angle of arrival of the signal(s)corresponds to or represents the direction from the reference deviceto the wearable device.
152 244 244 150 110 184 184 184 184 184 152 110 184 154 Optionally, the communication systemcan also include a range detector. The range detectoris configured to determine a distance from the reference deviceto the wearable devicebased on the signal(s). To illustrate, information indicative of attenuation of the signal(s)can be used to determine the distance. As one example, the attenuation of the signal(s)can be determined based on a received signal strength of the signal(s)and a transmitted signal strength of the signal(s). As another example, the communication systemcan determine the distance and direction to the wearable devicebased on phase and signal strength of the signal(s)as received at spatially diverse antennas of the antenna(s).
152 202 202 242 244 202 150 150 110 150 110 As another example, the communication systemcan optionally include a beamformer. The beamformercan include, be included within, or correspond to, or be coupled to the direction detector, the range detector, or both. In this example, the beamformermay be configured to sweep one or more beams about an area around the reference deviceto determine a direction from the reference deviceto the wearable device, a distance from the reference deviceto the wearable device, or both.
2 FIG. 2 FIG. 152 204 206 204 206 184 154 162 184 154 204 184 204 206 162 284 162 152 In, the communication systemalso includes a receiverand a modem. The receiverand the modemcan be configured to process the signal(s)received by the antenna(s)to extract the audio datafrom the signals. For example, the antenna(s)can provide the receiverwith electrical signals corresponding to the signals, and the receivercan amplify, filter, mix, or otherwise modify the electrical signals to generate a modulated signal. The modemcan demodulate the modulated signal to generate a bitstream representing the audio dataor the data packets. Depending on how the audio datais sent, the communication systemcan also include other components such as jitter buffers, analog-to-digital converters, etc., that are not shown infor the sake of simplicity.
162 184 210 190 160 210 160 242 244 In some embodiments, the audio dataextracted from (or determined based on) the signal(s)is stored at the memoryfor processing by the processor(s). In such embodiments, the directionality informationmay also be stored at the memory. The directionality informationcan include or correspond to direction information determined by the direction detector, range information determined by the range detector, or both.
140 180 162 160 162 140 190 240 162 140 240 162 162 The spatial audio generatoris configured to generate the spatial audio databased on the audio dataand directionality informationassociated with the audio data. In some embodiments, in addition to the spatial audio generator, the processor(s)can include a pre-processorthat is configured to modify the audio dataand provide the modified audio data to the spatial audio generator. For example, the pre-processorcan perform operations to emphasize some audio components of the audio data(e.g., target audio components), can perform operations to de-emphasize other audio components of the audio data(e.g., non-target audio components), or both.
2 FIG. 150 220 282 150 282 182 150 110 240 162 282 220 180 162 282 220 As an example, in, the reference deviceincludes one or more microphonesconfigured to capture soundin an area around the reference device. In this example the soundand the soundmay both include audio components representing non-speech sounds (e.g., noise) in an area that includes the reference deviceand the wearable device. The pre-processormay be configured to modify the audio databased on the soundcaptured at the microphone(s)to de-emphasize, in the spatial audio data, audio components that are present in both the audio dataand in the soundcaptured at the microphone(s).
150 180 150 222 150 190 180 180 2 FIG. Optionally, the reference devicecan include one or more sensors and the spatial audio datacan be processed in conjunction with data from the sensor(s) to form a multimedia stream that includes spatial audio. For example, in, the reference deviceincludes one or more camerasthat are configured to capture video data in an area around the reference device. In this example, the processor(s)may be configured to process the video data in conjunction with the spatial audio dataand to encode the video data and the spatial audio datafor communication to another device.
162 184 160 160 162 150 110 110 150 150 160 150 110 110 150 150 150 160 110 162 284 150 160 In the description above, the audio datais described as being transmitted via the same signal(s)that are used to determine the directionality information; however, in other embodiments, different signals are used to determine the directionality informationthan are used to transmit the audio data. For example, the reference devicecan transmit signals that are received by the wearable device. In this example, the wearable devicecan send information descriptive of the signals received from the reference deviceto the reference devicefor processing to determine the directionality information. To illustrate, the reference devicecan periodically or occasionally transmit beacon signals or advertisement packets that are received by the wearable device. In this illustrative example, the wearable devicecan send information such as angle of arrival, time of receipt, signal strength, etc., characteristic of the signals transmitted by the reference deviceto the reference device, and the reference devicecan determine the directionality informationbased on the information. As another example, the wearable devicecan transmit the audio datavia the data packetsand can also transmit other signals (e.g., advertisement packets, beacons signals, ranging and direction signals, etc.). In this example, the reference devicedetermines the directionality informationbased on the other signals.
3 FIG. 1 FIG. 3 FIG. 300 100 192 110 150 192 150 110 302 192 is a block diagram of a systemthat includes or corresponds to particular illustrative aspects of the systemofin accordance with some examples of the present disclosure. In the example illustrated in, the deviceis located remote from the wearable deviceand the reference device. For example, the devicecan communicate with the reference device, the wearable device, or both, via one or more networks. To illustrate, the devicecan include or correspond to one or more server computing devices or cloud computing devices.
3 FIG. 3 FIG. 110 120 182 110 252 254 252 184 150 252 162 192 302 In, the wearable deviceincludes the microphone(s)to capture the sound(e.g., speech of a user wearing the wearable device), the communication system, and the antenna(s). As described above, the communication systemis configured to exchange the signal(s)with the reference device. In the example of, the communication systemmay also be configured to send information, such as the audio datato the devicevia the network(s).
3 FIG. 2 FIG. 150 154 152 152 242 244 150 160 184 110 184 110 150 150 110 150 160 192 302 110 162 150 150 162 192 302 In, the reference deviceincludes the antenna(s)coupled to the communication system. The communication systemincludes the direction detector, and optionally, may also include the range detector, each of which operate as described with reference to. The reference deviceis configured to generate the directionality informationbased on the signalsexchanged with the wearable device. As explained above, exchange of the signalscan include transmissions from the wearable devicereceived by the reference device, transmissions from the reference devicereceived by the wearable device, or both. The reference deviceis configured to send the directionality informationto the devicevia the network(s). In some embodiments, the wearable devicesends the audio datato the reference device, and the reference devicesends the audio datato the devicevia the network(s).
3 FIG. 192 190 140 140 180 162 160 192 180 192 180 192 180 192 180 150 110 302 150 110 180 180 180 In, the deviceincludes the processor(s), which include the spatial audio generator. The spatial audio generatoris configured to generate the spatial audio databased on the audio dataand the directionality information. In some embodiments, the deviceis configured to provide the spatial audio dataas output to another device or to render the spatial audio data for consumption. For example, the devicecan include a spatial audio renderer that is configured to generate output sound based on the spatial audio data. As another example, the devicecan send the spatial audio datato another device for storage, for further transmission, for output, etc. To illustrate, the devicecan send the spatial audio datato the reference deviceor to the wearable devicevia the network(s). In this example, whichever of the reference deviceor the wearable devicereceives the spatial audio datacan store the spatial audio datato a memory for later consumption or transmission, or can forward the spatial audio datato another device.
110 182 162 162 160 192 180 180 110 150 180 110 150 180 110 150 180 As one example, during a call, the wearable devicecan capture the soundrepresenting speech of a user and generate the audio data. The audio dataand associated directionality informationcan be sent to the device(e.g., a server) to generate the spatial audio data. The spatial audio datacan be sent back to the wearable deviceor the reference devicefor transmission to one or more other participants in the call. In this example, the spatial audio datacan be rendered by devices associated with the other participant(s) of the call such that movement of the user of the wearable devicewithin a space that includes the reference deviceis rendered as movement of a sound source in the spatial audio data. To illustrate if the user of the wearable devicewalks around a room that includes the reference devicewhile speaking, the spatial audio dataprovided to the other participant(s) of the call reflects such movement.
4 FIG. 4 FIG. 400 192 402 190 190 140 402 406 404 404 162 160 404 162 184 190 160 184 190 242 244 402 408 410 180 depicts an embodimentof the deviceas an integrated circuitthat includes the one or more processors. The processor(s)ininclude the spatial audio generator. The integrated circuitalso includes a signal input, such as one or more bus interfaces, to enable input datato be received for processing. For example, the input datacan include the audio dataand the directionality information. As another example, the input datacan include the audio dataand information descriptive of the signal(s), and the processor(s)can determine the directionality informationbased on characteristics of the signal(s). To illustrate, the processor(s)can optionally include the direction detector, the range detector, or both. The integrated circuitalso includes a signal output, such as a bus interface, to enable sending of output data, such as the spatial audio data.
5 6 7 8 9 FIGS.,,,, and 5 9 FIGS.- 5 9 FIGS.- 5 9 FIGS.- are diagrams illustrating various non-limiting examples of devices operable to interact to generate spatial audio data in accordance with some examples of the present disclosure. The specific devices and combinations of devices illustrated inare merely illustrative and are not limiting. In other embodiments, a system to generate spatial audio data according to aspects disclosed herein can include one or more different devices than are illustrated in, a different arrangement of devices than illustrated in, or both.
5 FIG. 1 FIG. 1 4 FIGS.- 500 150 192 504 110 502 502 120 520 252 504 152 140 500 depicts a systemin which the reference deviceand the deviceofare integrated within a laptop computing deviceand the wearable devicecorresponds to a headset device. The headset deviceincludes the microphone(s), one or more speakers, and the communication system. The laptop computing deviceincludes the communication systemand the spatial audio generator. The systemmay optionally include other features of any of.
5 FIG. 120 502 252 152 184 184 120 184 504 502 520 In the example illustrated in, at least one of the microphone(s)is positioned to capture sound corresponding to or including speech of a user wearing the headset device. The communication systemand the communication systemare configured to exchange the signals. In some embodiments, the signalsinclude audio data representing the sound captured by the microphone(s). The signalscan also include other data, such as audio data sent by the laptop computing deviceto the headset devicefor output via the speaker(s), directionality information, etc.
152 502 504 152 160 152 184 184 184 502 504 504 242 244 140 180 120 2 FIG. 1 FIG. The communication systemis configured to determine information indicative of the position of the headset devicerelative to the laptop computing device. For example, the communication systemcan determine the directionality information. In some embodiments, the communication systemdetermines characteristics of the signalsthat are related to an angle of arrival of the signals, characteristics of the signalsthat are related to a range to the headset device, or both, and one or more other components of the laptop computing devicedetermine the directionality information. For example, in such embodiments the laptop computing devicemay also include the direction detector, the range detector, or both, of. The spatial audio generatoris configured to generate spatial audio data (e.g., the spatial audio dataof) based on audio data captured by the microphone(s)and the directionality information.
504 152 140 504 140 504 152 504 184 502 192 140 5 FIG. 3 FIG. Although the laptop computing deviceofis illustrated as including both the communication systemand the spatial audio generator, in other embodiments, the laptop computing devicedoes not include the spatial audio generator. For example, in some such embodiments, the laptop computing deviceincludes the communication system, and the laptop computing deviceis configured to send directionality information based on the signals, audio data from the headset device, or both, to a remote device, such as a server (e.g., the deviceof). In this example, the remote device can include the spatial audio generator.
502 504 120 500 As one example, during use, a user of the headset devicecan speak while moving relative to the laptop computing device. To illustrate, while on a call or a video conference, the user can move around a room. In this example, the spatial audio data can be generated such that a sound source corresponding to the user moves around in the spatial audio data, even though the user is not moving relative to the microphone(s). Thus, the systemenables generation of spatial audio data without complicated, special-purpose microphone arrays.
504 522 504 522 140 522 Optionally, the laptop computing devicecan include a camera. In embodiments in which the laptop computing deviceincludes the camera, the spatial audio data from the spatial audio generatorcan be combined with video data from the camerato generate a multimedia stream that includes the spatial audio data.
504 524 504 524 524 120 Optionally, the laptop computing devicecan include one or more additional microphones. In embodiments in which the laptop computing deviceincludes the microphone(s), audio data captured by the microphone(s)can be used to modify the audio data captured by the microphone(s), such as to de-emphasize noise components in the spatial audio data.
6 FIG. 1 FIG. 6 FIG. 600 150 192 604 110 602 602 622 602 622 602 120 620 252 depicts a systemin which the reference deviceand the deviceofare integrated within a game consoleand the wearable devicecorresponds to an extended reality (XR) headset device. In this context, extended reality includes virtual reality, augmented reality, mixed reality, or combinations thereof. The XR headset deviceis configured to be worn by a user such that a displayof the XR headset deviceis positioned in front of the user's eyes. In, in addition to the display, the XR headset deviceincludes the microphone(s), one or more speakers, and the communication system.
604 606 602 604 152 184 252 602 604 602 622 620 602 120 252 602 604 184 6 FIG. The game consoleis configured to communicate with a game controllerand with the XR headset device. To illustrate, in, the game consoleincludes the communication systemwhich is configured to exchange the signalswith the communication systemof the XR headset device. In some embodiments, the game consolecan send game information (e.g., control menus, in-game graphics, and game related audio) to the XR headset devicefor output to the user via the display, the speaker(s), or both. Additionally, the XR headset devicecan generate audio data based on sound (e.g., speech of the user) detected by the microphone(s). The communication systemof the XR headset devicecan send the audio data to the game consolevia the signals.
184 602 604 184 602 604 152 140 160 602 604 184 184 602 604 1 FIG. In some embodiments, multiple types of signalscan be exchanged by the XR headset deviceand the game console. For example, in addition to game information and audio data, the signalsexchanged by the XR headset deviceand the game consolecan include ranging and/or positioning signals used by the communication systemand/or the spatial audio generatorto determine a direction and/or distance (e.g., the directionality informationof) between the XR headset deviceand the game console. In some embodiments, a portion of the signalsincluding the game information and/or a portion of the signalsincluding the audio data is used to determine the direction and/or distance between the XR headset deviceand the game console.
152 604 602 604 152 160 152 184 184 184 602 604 604 242 244 140 180 120 2 FIG. 1 FIG. In some embodiments, the communication systemof the game consoleis configured to determine information indicative of the position of the XR headset devicerelative to the game console. For example, the communication systemcan determine the directionality information. In some embodiments, the communication systemdetermines characteristics of the signalsthat are related to an angle of arrival of the signals, characteristics of the signalsthat are related to a range to the XR headset device, or both, and one or more other components of the game consoledetermine the directionality information. For example, in such embodiments the game consolemay also include the direction detector, the range detector, or both, of. The spatial audio generatoris configured to generate spatial audio data (e.g., the spatial audio dataof) based on audio data captured by the microphone(s)and the directionality information.
604 152 140 604 140 604 152 604 184 602 192 140 6 FIG. 3 FIG. Although the game consoleofis illustrated as including both the communication systemand the spatial audio generator, in other embodiments, the game consoledoes not include the spatial audio generator. For example, in some such embodiments, the game consoleincludes the communication system, and the game consoleis configured to send directionality information based on the signals, audio data from the XR headset device, or both, to a remote device, such as a server (e.g., the deviceof). In this example, the remote device can include the spatial audio generator.
602 604 120 600 As one example, during use, a user of the XR headset devicecan speak while moving relative to the game console. To illustrate, while playing a multiplayer game, the user can move around a room and speak to other players. In this example, the other players can receive spatial audio data in which a sound source corresponding to the user moves around in the in-game audio, even though the user is not moving relative to the microphone(s). Thus, the systemenables generation of spatial audio data without complicated, special-purpose microphone arrays.
600 604 1 4 FIGS.- Optionally, the systemcan include other features of any of. For example, the game consolecan include a camera, one or more microphones, or a combination thereof.
7 FIG. 1 FIG. 1 4 FIGS.- 700 150 192 704 110 702 702 120 768 252 704 152 140 700 depicts a systemin which the reference deviceand the deviceofare integrated within a mobile communication device(e.g., a smart phone) and the wearable devicecorresponds to a pair of earbuds. At least one of the earbudsincludes the microphone(s), one or more speakers, and the communication system. The mobile communication deviceincludes the communication systemand the spatial audio generator. The systemmay optionally include other features of any of.
7 FIG. 7 FIG. 702 702 702 702 120 702 702 762 762 762 764 766 702 768 702 702 In the example illustrated in, the earbudsinclude an earbudA and an earbudB. The earbudA includes the microphone(s), which in this example can include a high signal-to-noise microphone positioned to capture the voice of a wearer of the earbudA. The earbudA can also include one or more other microphones configured to detect ambient sounds and spatially distributed to support beamforming, illustrated inas microphonesA,B, andC, an “inner” microphoneproximate to the wearer's ear canal (e.g., to assist with active noise cancelling), and a self-speech microphone, such as a bone conduction microphone configured to convert sound vibrations of the wearer's ear bone or skull into an audio signal. The earbudA also includes a speaker. The earbudB can be configured in a substantially similar manner as the earbudA.
7 FIG. 120 702 252 152 184 184 120 184 704 702 768 In the example illustrated in, at least one of the microphone(s)is positioned to capture sound corresponding to or including speech of a user wearing the earbudA. The communication systemand the communication systemare configured to exchange the signals. In some embodiments, the signalsinclude audio data representing the sound captured by the microphone(s). The signalscan also include other data, such as audio data sent by the mobile communication deviceto the earbudsfor output via the speaker, directionality information, etc.
152 702 704 152 160 152 184 184 184 702 704 704 242 244 140 180 120 1 FIG. 2 FIG. 1 FIG. The communication systemis configured to determine information indicative of the position of at least one of the earbudsrelative to the mobile communication device. For example, the communication systemcan determine the directionality informationof. In some embodiments, the communication systemdetermines characteristics of the signalsthat are related to an angle of arrival of the signals, characteristics of the signalsthat are related to a range to the earbud(s), or both, and one or more other components of the mobile communication devicedetermine the directionality information. For example, in such embodiments the mobile communication devicemay also include the direction detector, the range detector, or both, of. The spatial audio generatoris configured to generate spatial audio data (e.g., the spatial audio dataof) based on audio data captured by the microphone(s)and the directionality information.
704 152 140 704 140 704 152 704 184 702 192 140 7 FIG. 3 FIG. Although the mobile communication deviceofis illustrated as including both the communication systemand the spatial audio generator, in other embodiments, the mobile communication devicedoes not include the spatial audio generator. For example, in some such embodiments, the mobile communication deviceincludes the communication system, and the mobile communication deviceis configured to send directionality information based on the signals, audio data from the earbuds, or both, to a remote device, such as a server (e.g., the deviceof). In this example, the remote device can include the spatial audio generator.
702 704 120 700 As one example, during use, a user of the earbudscan speak while moving relative to the mobile communication device. To illustrate, while on a call or a video conference, the user can move around a room. In this example, the spatial audio data can be generated such that a sound source corresponding to the user moves around in the spatial audio data, even though the user is not moving relative to the microphone(s). Thus, the systemenables generation of spatial audio data without complicated, special-purpose microphone arrays.
704 722 704 722 140 722 Optionally, the mobile communication devicecan include a camera. In embodiments in which the mobile communication deviceincludes the camera, the spatial audio data from the spatial audio generatorcan be combined with video data from the camerato generate a multimedia stream that includes spatial audio data.
704 720 704 720 720 120 Optionally, the mobile communication devicecan include one or more additional microphones. In embodiments in which the mobile communication deviceincludes the microphone(s), audio data captured by the microphone(s)can be used to modify the audio data captured by the microphone(s), such as to de-emphasize noise components in the spatial audio data.
700 702 700 702 184 704 702 702 252 702 700 7 FIG. Although the systemillustrates two earbudsof a pair, in other embodiments, the systemcan include a single earbud (e.g., the earbudA) that is configured to exchange the signalswith the mobile communication device. For example, the user can wear both earbuds, but only one of the earbudsincludes the communication system. Further, althoughillustrates earbudsconfigured to be worn with a portion that extends into or covers the user's ear canal, in other embodiments, the systemincludes other in-ear or over-ear devices.
8 FIG. 1 FIG. 1 4 FIGS.- 800 150 192 804 110 802 802 120 252 802 854 856 856 804 152 140 700 depicts a systemin which the reference deviceand the deviceofare integrated within a vehicle(e.g., a piloted or unpiloted aerial vehicle, a car, a train, a bus, a ship or boat, etc.) and the wearable devicecorresponds to XR glasses(e.g., augmented-reality, mixed-reality, or virtual-reality glasses). XR glassesincludes the microphone(s)and the communication system. The XR glassesinclude a holographic projection unitconfigured to project visual data onto a surface of a lensor to reflect the visual data off of a surface of the lensand onto the wearer's retina. The vehicleincludes the communication systemand the spatial audio generator. The systemmay optionally include other features of any of.
8 FIG. 120 802 252 152 184 184 120 184 804 802 854 802 In the example illustrated in, at least one of the microphone(s)is positioned to capture sound corresponding to or including speech of a user wearing the XR glasses. The communication systemand the communication systemare configured to exchange the signals. In some embodiments, the signalsinclude audio data representing the sound captured by the microphone(s). The signalscan also include other data, such as data sent by the vehicleto the XR glassesfor output to the user (e.g., via the holographic projection unitor one or more speakers of the XR glasses).
152 802 804 152 160 152 184 184 184 802 804 804 242 244 140 180 120 2 FIG. 1 FIG. The communication systemis configured to determine information indicative of the position of the XR glassesrelative to the vehicle. For example, the communication systemcan determine the directionality information. In some embodiments, the communication systemdetermines characteristics of the signalsthat are related to an angle of arrival of the signals, characteristics of the signalsthat are related to a range to the XR glasses, or both, and one or more other components of the vehicledetermine the directionality information. For example, in such embodiments the vehiclemay also include the direction detector, the range detector, or both, of. The spatial audio generatoris configured to generate spatial audio data (e.g., the spatial audio dataof) based on audio data captured by the microphone(s)and the directionality information.
804 152 140 804 140 804 152 804 184 802 192 140 8 FIG. 3 FIG. Although the vehicleofis illustrated as including both the communication systemand the spatial audio generator, in other embodiments, the vehicledoes not include the spatial audio generator. For example, in some such embodiments, the vehicleincludes the communication system, and the vehicleis configured to send directionality information based on the signals, audio data from the XR glasses, or both, to a remote device, such as a server (e.g., the deviceof). In this example, the remote device can include the spatial audio generator.
802 804 120 800 As one example, during use, a user of the XR glassescan speak while moving relative to the vehicle. To illustrate, while on a call or a video conference, the user can move around a scene. In this example, the spatial audio data can be generated such that a sound source corresponding to the user moves around in the spatial audio data, even though the user is not moving relative to the microphone(s). Thus, the systemenables generation of spatial audio data without complicated, special-purpose microphone arrays.
804 802 804 152 804 804 802 804 804 804 804 802 804 804 804 804 804 804 802 In some embodiments, the vehicleand the user of the XR glassesmove together. For example, the user can be a passenger of the vehicle. In such embodiments, the communication systemcan be disposed on or in a portion of the vehiclethat the user moves with respect to. To illustrate, when the vehicleis an aircraft, the user can move along an aisle of the aircraft. In some embodiments, the user of the XR glassescan remain stationary and the vehiclecan move relative to the user. For example, the vehiclecan include an unmanned aerial vehicle (UAV) that is configured to move around a scene or around the user. In this example, the movement of the vehiclerelative to the user can be reflected in directionality information indicating the distance and/or direction from the vehicleto the XR glasses. In this example, the spatial audio data represents a sound source associated with the user moving as the vehiclemoves. Alternatively, in some embodiments, the vehicleincludes an onboard positioning system (e.g., a local positioning system, a global positioning system, or both) configured to detect movement of the vehicle. In such embodiments, the movement of the vehiclecan be removed from the directionality information such that a sound source corresponding to the user is stationary in the spatial audio data when the vehiclemoves and the sound source moves when the user moves in a manner that changes the distance and/or direction from the vehicleto the XR glasses.
804 822 804 822 140 822 Optionally, the vehiclecan include a camera. In embodiments in which the vehicleincludes the camera, the spatial audio data from the spatial audio generatorcan be combined with video data from the camerato generate a multimedia stream that includes spatial audio data.
804 820 804 820 820 120 Optionally, the vehiclecan include one or more additional microphones. In embodiments in which the vehicleincludes the microphone(s), audio data captured by the microphone(s)can be used to modify the audio data captured by the microphone(s), such as to de-emphasize noise components in the spatial audio data.
9 FIG. 1 FIG. 900 150 192 904 110 902 902 120 252 120 112 120 252 252 120 252 depicts a systemin which the reference deviceand the deviceofare integrated within a camera, and the wearable devicecorresponds to a portable microphone device. The portable microphone deviceincludes the microphone(s)and the communication system. In the particular example illustrated, at least one of the microphone(s)is configured to be worn on the chest of the user(e.g., as a lapel microphone), but in other embodiments, the microphone(s)can be positioned elsewhere on the user's body or clothing. Further, in the example illustrated, the communication systemis disposed within a belt-mounted unit; however, in other embodiments, the communication systemand the microphone(s)are disposed within a single housing, or the communication systemis configured to be worn elsewhere on the user's body or clothing.
904 152 140 900 904 920 1 4 FIGS.- 9 FIG. The cameraincludes the communication systemand the spatial audio generator. The systemmay optionally include other features of any of. For example, in, the cameraincludes one or more microphones.
9 FIG. 120 112 252 152 184 184 120 184 In the example illustrated in, at least one of the microphone(s)is positioned to capture sound corresponding to or including speech of the user. The communication systemand the communication systemare configured to exchange the signals. In some embodiments, the signalsinclude audio data representing the sound captured by the microphone(s). The signalscan also include other data.
152 902 904 152 160 152 184 184 184 902 904 904 242 244 140 180 120 2 FIG. 1 FIG. The communication systemis configured to determine information indicative of the position of the portable microphone devicerelative to the camera. For example, the communication systemcan determine the directionality information. In some embodiments, the communication systemdetermines characteristics of the signalsthat are related to an angle of arrival of the signals, characteristics of the signalsthat are related to a range to the portable microphone device, or both, and one or more other components of the cameradetermine the directionality information. For example, in such embodiments the cameramay also include the direction detector, the range detector, or both, of. The spatial audio generatoris configured to generate spatial audio data (e.g., the spatial audio dataof) based on audio data captured by the microphone(s)and the directionality information.
904 152 140 904 140 904 152 904 184 902 192 140 9 FIG. 3 FIG. Although the cameraofis illustrated as including both the communication systemand the spatial audio generator, in other embodiments, the cameradoes not include the spatial audio generator. For example, in some such embodiments, the cameraincludes the communication system, and the camerais configured to send directionality information based on the signals, audio data from the portable microphone device, or both, to a remote device, such as a server (e.g., the deviceof). In this example, the remote device can include the spatial audio generator.
902 904 112 112 120 900 140 904 904 920 920 120 As one example, during use, a user of the portable microphone devicecan speak while moving relative to the camera. To illustrate, while recording a video segment, the user can move around a scene. In this example, the spatial audio data can be generated such that a sound source corresponding to the usermoves around in the spatial audio data, even though the useris not moving relative to the microphone(s). Thus, the systemenables generation of spatial audio data without complicated, special-purpose microphone arrays. Optionally, the spatial audio data from the spatial audio generatorcan be combined with video data from the camerato generate a multimedia stream that includes spatial audio data. In embodiments in which the cameraincludes the optional microphone(s), audio data captured by the microphone(s)can be used to modify the audio data captured by the microphone(s), such as to de-emphasize noise components in the spatial audio data.
5 9 FIGS.- 5 9 FIGS.- 5 9 FIGS.- 1 FIG. 110 502 602 702 802 902 110 150 504 604 704 804 904 150 150 192 The examples illustrated inare merely intended to highlight specific use cases and configurations, and are not limiting. In particular, the wearable devicecan include, correspond to, or be included within devices other than the headset device, the XR headset device, the earbuds, the XR glasses, and the portable microphone deviceof. In other non-limiting examples, the wearable devicecan include, correspond to, or be included in devices such as a pendant or broach, a smart watch, a helmet, or an article of smart clothing. Likewise, the reference devicecan include, correspond to, or be included within devices other than the laptop computing device, the game console, the mobile communication device, the vehicle, and the cameraof. In other non-limiting examples, the reference devicecan include, correspond to, or be included in devices such as a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a desktop computer, a tablet computer, a personal digital assistant device, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of-things (IoT) device, a base station, or any combination thereof. Further, as explained above, the reference deviceand the deviceofcan be integrated within a single device or distributed among two or more distinct devices that communicate with one another (e.g., via one or more local or wide area networks).
10 FIG. 1 FIG. 1000 1000 110 150 192 152 190 140 100 1000 1000 Referring to, a particular embodiment of a methodof generating spatial audio data is shown. In a particular aspect, one or more operations of the methodare performed by at least one of the wearable device, the reference device, the device, the communication system, the processor(s), the spatial audio generator, the systemof, or a combination thereof. One technical benefit of the methodis that the methodenables generation of spatial audio data without complicated, special-purpose microphone arrays.
1000 1002 110 162 150 192 162 182 120 110 1 FIG. In a particular aspect, the methodincludes, at block, obtaining, at one or more processors, audio data captured by a microphone of a wearable device. For example, a user can wear the wearable device in a manner that positions the microphone to capture speech or other sounds produced by the user. In this example, a communication system of the wearable device can send audio data representing the sound captured by the microphone to the reference device or to another device. To illustrate, the wearable deviceofcan send the audio datato the reference device, to the device, or both. In this example, the audio datarepresents the soundcaptured by the microphone(s)of the wearable device.
1000 1004 152 150 160 184 1 FIG. The methodalso includes, at block, determining, by the one or more processors, directionality information indicative of a direction between the microphone and the reference device based on one or more signals exchanged between the wearable device and the reference device. For example, the communication systemof the reference deviceofcan determine the directionality informationbased on the signals. The signals used determine the directionality information can include signals that encode the audio data or other signals, such as beacon signals or signals including advertisement packets. In some embodiments, the signals are transmitted in accordance with a data communication protocol, such as a BLUETOOTH® communication protocol or a WIFI® communication protocol.
1000 1006 1000 1000 The methodincludes, at block, generating, at the one or more processors, spatial audio data (e.g., ambisonics data, such as first-order or higher-order ambisonics data) based on the audio data and the directionality information. For example, the directionality information can be used to specify a location of a sound source corresponding to the audio data. In this example, the sound source moves within the spatial audio data as the wearable device moves relative to the reference device. To illustrate, the microphone can be configured to capture sound corresponding to the audio data at a fixed location relative to a source of sound (e.g., the user's mouth). In this illustrative example, at a second time, the user can move to a different location, resulting in movement of the wearable device relative to the reference device. In this situation, the methodcan include determining updated directionality information based on one or more additional signals exchanged between the wearable device and the reference device. The methodcan also include generating updated spatial audio data based on the updated directionality information. In the updated spatial audio data movement over time of the wearable device relative to the reference device is represented as movement of the source of the sound.
1000 In some embodiments, the one or more signals exchanged between the wearable device and the reference device include signals encoding the audio data. For example, the signal(s) can include encoded data representing the audio data (e.g., in a set of data packets), and the methodcan include decoding the encoded data to generate the audio data.
Determining the directionality information can include, for example, determining an angle of arrival of the one or more signals. In this example, the angle of arrival corresponds to or indicates a direction from the reference device to the wearable device. Determining the directionality information can also, or alternatively, include determining range information indicative of a distance (or a change of distance) between the microphone and the reference device. To illustrate, the range information can be determined based on a signal strength indicator associated with the signals. For example, a transmitted signal strength of the signals can be compared to a received signal strength of the signals to estimate a distance traversed by the signals.
1000 150 222 162 2 FIG. In some embodiments, the methodcan also include obtaining video data associated with the audio data, processing the video data in conjunction with the spatial audio data, and encoding the video data and the spatial audio data for transmission or storage. For example, the reference deviceofincludes the camera(s)which can capture video data associated with the audio data. In this example, the video data and the spatial audio data can be combined to generate a multimedia file or a multimedia stream.
150 220 282 150 162 162 282 220 150 2 FIG. In some embodiments, the audio data can be modified before the spatial audio data is generated. For example, the audio data can be modified to emphasize target audio components (e.g., speech), to de-emphasize non-target audio components (e.g., noise), or both. As one example, the reference deviceofincludes the microphone(s), which are configured to capture soundnear the reference device. In this example, the audio datacan be processed to de-emphasize, in the spatial audio data, audio components that are present in both the audio dataand in the soundcaptured at microphone(s)of the reference device.
1000 1000 10 FIG. 10 FIG. 11 FIG. The methodofmay be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit such as a central processing unit (CPU), a digital signal processor (DSP), a controller, another hardware device, firmware device, or any combination thereof. As an example, the methodofmay be performed by a processor that executes instructions, such as described with reference to.
11 FIG. 11 FIG. 1 FIG. 1 10 FIGS.- 1100 1100 1100 150 192 1100 Referring to, a block diagram of a particular illustrative embodiment of a device is depicted and generally designated. In various embodiments, the devicemay have more or fewer components than illustrated in. In an illustrative embodiment, the devicecorresponds to or includes the reference device, the device, or both, of. In an illustrative embodiment, the devicemay perform one or more operations described with reference to.
1100 1106 1100 1110 190 1106 1110 1110 1108 1136 1138 140 1 FIG. In a particular embodiment, the deviceincludes a processor(e.g., a CPU). The devicemay include one or more additional processors(e.g., one or more DSPs). In a particular aspect, the processor(s)ofcorrespond to the processor, the processors, or a combination thereof. The processorsmay include a speech and music coder-decoder (CODEC)that includes a voice coder (“vocoder”) encoder, a vocoder decoder, the spatial audio generator, or a combination thereof.
1100 210 1134 210 212 1110 1106 140 1100 206 1150 154 The devicemay include the memoryand a CODEC. The memorymay include the instructions, that are executable by the one or more additional processors(or the processor) to implement the functionality described with reference to the spatial audio generator. The devicemay include the modemcoupled, via a transceiver, to the antenna(s).
1100 1128 1126 1192 220 1134 1134 1102 1104 1134 220 1104 1108 1108 140 110 140 1108 1134 1134 1102 1192 The devicemay include a displaycoupled to a display controller. One or more speakers, one or more microphones, or a combination thereof, which may be coupled to the CODEC. The CODECmay include a digital-to-analog converter (DAC), an analog-to-digital converter (ADC), or both. In a particular embodiment, the CODECmay receive analog signals from the microphone(s), convert the analog signals to digital signals using the analog-to-digital converter, and provide the digital signals to the speech and music codec. The speech and music codecmay process the digital signals, and the digital signals may further be processed by the spatial audio generator. For example, audio components present in the digital signals can be subtracted from audio data received from the wearable deviceto de-emphasize such audio components from spatial audio data generated by the spatial audio generator. In a particular embodiment, the speech and music codecmay provide digital signals to the CODEC. The CODECmay convert the digital signals to analog signals using the digital-to-analog converterand may provide the analog signals to the speaker(s).
1100 1122 210 1106 1110 1126 1134 206 1122 1130 1144 1122 1128 1130 1192 220 222 154 1144 1122 1128 1130 1192 220 222 154 1144 1122 11 FIG. In a particular embodiment, the devicemay be included in a system-in-package or system-on-chip device. In a particular embodiment, the memory, the processor, the processors, the display controller, the CODEC, and the modemare included in the system-in-package or system-on-chip device. In a particular embodiment, an input deviceand a power supplyare coupled to the system-in-package or the system-on-chip device. Moreover, in a particular embodiment, as illustrated in, the display, the input device, the speaker(s), the microphone(s), one or more cameras, the antenna(s), and the power supplyare external to the system-in-package or the system-on-chip device. In a particular embodiment, each of the display, the input device, the speaker(s), the microphone(s), the camera(s), the antenna(s), and the power supplymay be coupled to a component of the system-in-package or the system-on-chip device, such as an interface or a controller.
1122 1150 206 154 1150 206 152 184 110 140 180 162 110 160 184 1 FIG. In a particular aspect, the system-in-package or system-on-chip devicealso includes a transceivercoupled to the modemand the antenna(s). The transceiver, the modem, and possibly other component, correspond to the communication systemof. The communication system is operable to exchange the signalswith the wearable device, and the spatial audio generatoris operable to generate spatial audio databased on audio datafrom the wearable deviceand directionality informationdetermined based on the signals.
1100 The devicemay include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.
154 152 150 140 190 192 240 204 206 202 242 244 In conjunction with the described embodiments, an apparatus includes means for obtaining audio data captured by a microphone of a wearable device. For example, the means for obtaining audio data captured by the microphone of the wearable device can include or correspond to the antenna(s), the communication system, the reference device, the spatial audio generator, the processor(s), the device, the pre-processor, the receiver, the modem, the beamformer, the direction detector, the range detector, one or more other circuits or components configured to obtain audio data captured by a microphone of a wearable device, or any combination thereof.
152 150 140 190 192 204 202 242 244 The apparatus also includes means for determining directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device. For example, the means for determining directionality information can include or correspond to the communication system, the reference device, the spatial audio generator, the processor(s), the device, the receiver, the beamformer, the direction detector, the range detector, one or more other circuits or components configured to determine directionality information based on exchanged signals, or any combination thereof.
150 140 190 192 The apparatus also includes means for generating spatial audio data based on the audio data and the directionality information. For example, the means for generating the spatial audio data can include or correspond to the reference device, the spatial audio generator, the processor(s), the device, one or more other circuits or components configured to generate the spatial audio data, or any combination thereof.
210 212 190 1110 1106 162 120 110 160 150 184 180 In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as the memory) includes instructions (e.g., the instructions) that, when executed by one or more processors (e.g., the processor(s), the processor(s), or the processor), cause the one or more processors to obtain audio data (e.g., the audio data) captured by a microphone (e.g., the microphone(s)) of a wearable device (e.g., the wearable device), determine directionality information (e.g., the directionality information) indicative of a direction between the microphone and a reference device (e.g., the reference device) based on one or more signals (e.g., the signals) exchanged between the wearable device and the reference device, and generate spatial audio data (e.g., the spatial audio data) based on the audio data and the directionality information.
Particular aspects of the disclosure are described below in sets of interrelated Examples:
According to Example 1, a device includes a memory configured to store audio data and one or more processors configured to: obtain the audio data captured by a microphone of a wearable device; determine, based on one or more signals exchanged between the wearable device and a reference device, directionality information indicative of a direction of the microphone relative to the reference device; and process the audio data based on the directionality information to generate spatial audio data.
Example 2 includes the device of Example 1, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and wherein the one or more processors are configured to update the spatial audio data over time to represent movement of the wearable device relative to the reference device as movement of the source of the sound.
Example 3 includes the device of Example 1 or Example 2, wherein the one or more signals include encoded data, wherein the one or more processors are configured to receive the one or more signals and decode the encoded data to generate the audio data.
Example 4 includes the device of any of Examples 1 to 3, wherein the one or more processors are configured to obtain the audio data from one or more data packets of the one or more signals.
Example 5 includes the device of any of Examples 1 to 4, wherein the one or more processors are further configured to determine, based on a signal strength indicator associated with the one or more signals, range information indicative of a distance between the microphone and the reference device.
Example 6 includes the device of any of Examples 1 to 5, wherein the one or more processors are further configured to determine, based on a signal strength indicator associated with the one or more signals, a change of distance between the microphone and the reference device.
Example 7 includes the device of any of Examples 1 to 6, wherein the one or more processors are configured to determine the directionality information based on an angle of arrival of the one or more signals.
Example 8 includes the device of any of Examples 1 to 7 and further includes one or more antennas configured to transmit a signal of the one or more signals, to receive a signal of the one or more signals, or both.
Example 9 includes the device of any of Examples 1 to 8, wherein the one or more processors are further configured to determine, based on a received signal strength of the one or more signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
Example 10 includes the device of any of Examples 1 to 9, wherein the spatial audio data includes ambisonics data.
Example 11 includes the device of any of Examples 1 to 10 and further includes a camera coupled to the one or more processors and configured to capture video data, wherein the one or more processors are configured to process the video data in conjunction with the spatial audio data and to encode the video data and the spatial audio data for communication to another device.
Example 12 includes the device of any of Examples 1 to 11 and further includes a second microphone coupled to the one or more processors, wherein the one or more processors are configured to modify the audio data based on sound captured at the second microphone.
Example 13 includes the device of Example 12, wherein the one or more processors are configured to modify the audio data to de-emphasize, in the spatial audio data, audio components that are present in both the audio data and in the sound captured at the second microphone.
Example 14 includes the device of any of Examples 1 to 13, wherein the one or more processors and the memory are integrated within the reference device.
Example 15 includes the device of any of Examples 1 to 13, wherein the one or more processors and the memory are integrated within the wearable device.
Example 16 includes the device of any of Examples 1 to 15, wherein the one or more processors and the memory are integrated into at least one of a smart speaker, a speaker bar, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a tuner, a camera, a navigation device, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a communication device, an internet-of-things (IoT) device, an extended reality (XR) device, a base station, or a mobile device.
Example 17 includes the device of any of Examples 1 to 16, wherein the wearable device corresponds to or includes a headset device or one or more earbuds.
According to Example 18, a method includes obtaining, at one or more processors, audio data captured by a microphone of a wearable device; determining, by the one or more processors, directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device; and generating, at the one or more processors, spatial audio data based on the audio data and the directionality information.
Example 19 includes the method of Example 18, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and further including, after determining the directionality information, determining updated directionality information; and generating updated spatial audio data, wherein the updated spatial audio data represents movement over time of the wearable device relative to the reference device as movement of the source of the sound.
Example 20 includes the method of Example 18 or Example 19, wherein the one or more signals include signals encoding the audio data.
Example 21 includes the method of any of Examples 18 to 20, wherein the one or more signals include encoded data, and further comprising receiving the one or more signals and decoding the encoded data to generate the audio data.
Example 22 includes the method of any of Examples 18 to 21 and further includes determining, based on a signal strength indicator, range information indicative of a distance between the microphone and the reference device.
Example 23 includes the method of any of Examples 18 to 22 and further includes determining, based on a signal strength indicator, a change of distance between the microphone and the reference device.
Example 24 includes the method of any of Examples 18 to 23, wherein the directionality information is based on an angle of arrival of the one or more signals.
Example 25 includes the method of any of Examples 18 to 24, wherein the one or more signals are transmitted in accordance with a BLUETOOTH® communication protocol.
Example 26 includes the method of any of Examples 18 to 25 and further includes determining, based on a received signal strength of the one or more signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
Example 27 includes the method of any of Examples 18 to 26, wherein the spatial audio data includes ambisonics data.
Example 28 includes the method of any of Examples 18 to 27 and further includes obtaining video data associated with the audio data; processing the video data in conjunction with the spatial audio data; and encoding the video data and the spatial audio data for transmission or storage.
Example 29 includes the method of any of Examples 18 to 28 and further includes modifying the audio data based on sound captured at a second microphone.
Example 30 includes the method of Example 29 and further includes modifying the audio data to de-emphasize, in the spatial audio data, audio components that are present in the audio data and in the sound captured at a microphone of the reference device.
According to Example 31, a non-transitory computer-readable device storing instructions that are executable by one or more processors to cause the one or more processors to obtain audio data captured by a microphone of a wearable device; determine directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device; and generate spatial audio data based on the audio data and the directionality information.
Example 32 includes the non-transitory computer-readable device of Example 31, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and wherein the instructions are further executable to: after determining the directionality information, determine updated directionality information; and generate updated spatial audio data based on the updated directionality information, wherein the updated spatial audio data represents movement over time of the wearable device relative to the reference device as movement of the source of the sound.
Example 33 includes the non-transitory computer-readable device of Example 31 or Example 32, wherein the one or more signals include encoded data, wherein the instructions are further executable to decode the encoded data to generate the audio data.
Example 34 includes the non-transitory computer-readable device of any of Examples 31 to 33, wherein the audio data is obtained from one or more data packets of the one or more signals.
Example 35 includes the non-transitory computer-readable device of any of Examples 31 to 34, wherein the instructions are further executable to determine, based on a signal strength indicator, range information indicative of a distance between the microphone and the reference device.
Example 36 includes the non-transitory computer-readable device of any of Examples 31 to 35, wherein the instructions are further executable to determine, based on a signal strength indicator, a change of distance between the microphone and the reference device.
Example 37 includes the non-transitory computer-readable device of any of Examples 31 to 36, wherein the directionality information is based on an angle of arrival of the one or more signals.
Example 38 includes the non-transitory computer-readable device of any of Examples 31 to 37, wherein the one or more signals are transmitted in accordance with a BLUETOOTH® communication protocol.
Example 39 includes the non-transitory computer-readable device of any of Examples 31 to 38, wherein the instructions are further executable to determine, based on a received signal strength of the one or more signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
Example 40 includes the non-transitory computer-readable device of any of Examples 31 to 39, wherein the spatial audio data includes ambisonics data.
Example 41 includes the non-transitory computer-readable device of any of Examples 31 to 40, wherein the instructions are further executable to obtain video data associated with the audio data; process the video data in conjunction with the spatial audio data; and encode the video data and the spatial audio data for transmission or storage.
Example 42 includes the non-transitory computer-readable device of any of Examples 31 to 41, wherein the instructions are further executable to modify the audio data based on sound captured at a second microphone.
Example 43 includes the non-transitory computer-readable device of Example 42, wherein the instructions are further executable to modify the audio data to de-emphasize, in the spatial audio data, audio components that are present in the audio data and in the sound captured at a microphone of the reference device.
According to Example 44, an apparatus includes means for obtaining audio data captured by a microphone of a wearable device; means for determining directionality information indicative of a direction between the microphone and a reference device based on one or more signals exchanged between the wearable device and the reference device; and means for generating spatial audio data based on the audio data and the directionality information.
Example 45 includes the apparatus of Example 44, wherein the microphone captures the audio data at a fixed location relative to a source of sound, and further includes means for determining updated directionality information after determining the directionality information; and means for generating updated spatial audio data based on the updated directionality information, wherein the updated spatial audio data represents movement over time of the wearable device relative to the reference device as movement of the source of the sound.
Example 46 includes the apparatus of Example 44 or Example 45, wherein the one or more signals include an encoded version of the audio data.
Example 47 includes the apparatus of any of Examples 44 to 46, wherein the audio data is obtained from one or more data packets of the one or more signals.
Example 48 includes the apparatus of any of Examples 44 to 47 and further includes means for determining, based on a signal strength indicator, range information indicative of a distance between the microphone and the reference device.
Example 49 includes the apparatus of any of Examples 44 to 48 and further includes means for determining, based on a signal strength indicator, a change of distance between the microphone and the reference device.
Example 50 includes the apparatus of any of Examples 44 to 49, wherein the directionality information is based on an angle of arrival of the one or more signals.
Example 51 includes the apparatus of any of Examples 44 to 50, wherein the one or more signals are transmitted in accordance with a BLUETOOTH® communication protocol.
Example 52 includes the apparatus of any of Examples 44 to 51 and further includes means for determining, based on a received signal strength of the one or more signals, range information associated with a distance between the microphone and the reference device, wherein the audio data is processed further based on the range information to generate the spatial audio data.
Example 53 includes the apparatus of any of Examples 44 to 52, wherein the spatial audio data includes ambisonics data.
Example 54 includes the apparatus of any of Examples 44 to 53 and further includes: means for obtaining video data associated with the audio data; means for processing the video data in conjunction with the spatial audio data; and means for encoding the video data and the spatial audio data for transmission or storage.
Example 55 includes the apparatus of any of Examples 44 to 54 and further includes means for modifying the audio data based on sound captured at a second microphone.
Example 56 includes the apparatus of Example 55 and further includes means for modifying the audio data to de-emphasize, in the spatial audio data, audio components that are present in the audio data and in the sound captured at a microphone of the reference device.
Those of skill would further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor executable instructions depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, such embodiment decisions are not to be interpreted as causing a departure from the scope of the present disclosure.
The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor may read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.
The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
August 16, 2023
August 11, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.