Patentable/Patents/US-20260188337-A1
US-20260188337-A1

Enhanced Audio File Generator

PublishedJuly 2, 2026
Assigneenot available in USPTO data we have
Technical Abstract

This disclosure is directed to an enhanced audio file generator. One aspect is a method of enhancing input speech in an input audio file, the method comprising receiving the input audio file representing the input speech, wherein the input audio file is recorded at an audio recording device, and generating an enhanced audio file by applying an audio transformation model to the input audio file, wherein applying the audio transformation model to generate the enhanced audio file comprises extracting parameters defining audio features from the input audio file, the parameters including a noise parameter defining noise in the input audio file and one or more other preset parameters respectively defining other audio features, synthesizing clean speech based on the extracted parameters including the noise parameter, wherein synthesizing the clean speech comprises transforming the noise parameter to defined value(s); and generating the enhanced audio file with the synthesized clean speech.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

receiving an input audio file including input speech; determining parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other non-noise parameters respectively defining other audio features; transforming the noise parameter to at least one defined value that renders the noise inaudible; and synthesizing, in an enhanced audio file, a clean version of the input speech based on the parameters. . A computer-implemented method comprising:

2

claim 1 . The computer-implemented method of, wherein receiving the input audio file including the input speech comprises recording the input audio file by way of a microphone of an audio recording device.

3

claim 2 . The computer-implemented method of, wherein at least one of the receiving, determining, transforming, or synthesizing steps are performed in real-time as a user records audio with the audio recording device.

4

claim 1 . The computer-implemented method of, wherein synthesizing the clean version of the input speech based on the parameters comprises synthesizing the clean version of the input speech by way of a trained neural network.

5

claim 4 providing a training audio file representing training speech; extracting training parameters from the training audio file; synthesizing, by way of the trained neural network, reconstructed speech based on the parameters; generating an output audio file with the reconstructed speech; and comparing the output audio file to the training audio file. . The computer-implemented method of, wherein the trained neural network was trained by:

6

claim 5 . The computer-implemented method of, wherein during the training of the trained neural network the noise parameter prior to transformation is incorporated into the output audio file prior to comparing the output audio file to the training audio file.

7

claim 4 . The computer-implemented method of, wherein the trained neural network is trained using a combination of components corresponding to the noise parameter and the one or more other non-noise parameters.

8

claim 1 . The computer-implemented method of, wherein the at least one defined value is zero.

9

claim 1 . The computer-implemented method of, wherein the clean version of the input speech is synthesized without referencing the input audio file.

10

claim 1 . The computer-implemented method of, wherein the one or more other non-noise parameters comprise one or more of: phonemes, pitch salience, voice timbre, or voice volume.

11

claim 1 . The computer-implemented method of, wherein the parameters also include recording environment data, wherein the recording environment data includes information regarding how or where the input audio file was recorded, and wherein synthesizing the clean version of the input speech based on the parameters comprises synthesizing the clean version of the input speech based at least in part on the recording environment data.

12

claim 11 . The computer-implemented method of, wherein the recording environment data is used to remove, from the input speech, abnormalities captured on specific recording devices.

13

claim 1 . The computer-implemented method of, wherein synthesizing the clean version of the input speech based on the parameters comprises removing hard consonant sounds from the input speech.

14

claim 1 . The computer-implemented method of, wherein synthesizing the clean version of the input speech based on the parameters comprises relacing a voice timbre parameter of the parameters with a different voice timbre parameter from a different voice recording.

15

claim 1 uploading the clean version of the input speech to a media delivery system. . The computer-implemented method of, further comprising:

16

claim 1 adding background music to the clean version of the input speech. . The computer-implemented method of, further comprising:

17

receiving an input audio file including input speech; determining parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other non-noise parameters respectively defining other audio features; transforming the noise parameter to at least one defined value that renders the noise inaudible; and synthesizing, in an enhanced audio file, a clean version of the input speech based on the parameters. . A non-transitory computer-readable medium, storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations comprising:

18

claim 17 . The non-transitory computer-readable medium of, wherein the clean version of the input speech is synthesized without referencing the input audio file.

19

claim 17 . The non-transitory computer-readable medium of, wherein the parameters also include recording environment data, wherein the recording environment data includes information regarding how or where the input audio file was recorded, and wherein synthesizing the clean version of the input speech based on the parameters comprises synthesizing the clean version of the input speech based at least in part on the recording environment data.

20

one or more processors; memory; and receiving an input audio file including input speech; determining parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other non-noise parameters respectively defining other audio features; transforming the noise parameter to at least one defined value that renders the noise inaudible; and synthesizing, in an enhanced audio file, a clean version of the input speech based on the parameters. program instructions, stored in the memory, that upon execution by the one or more processors cause the computing system to perform operations comprising: . A computing system comprising:

Detailed Description

Complete technical specification and implementation details from the patent document.

This application is a continuation of and claims priority to U.S. patent application Ser. No. 17/934,906, filed Sep. 23, 2022, which is hereby incorporated by reference in its entirety.

In order to produce high quality voice recordings, a professional studio with professional audio equipment is generally needed. The studio can be sound proofed to reduce background noise. The audio equipment can include a professional quality microphone, a pop filter, a multi-channel recorder, audio mixing and equalizing hardware, a good computer, headphones, etc. Many people do not have access to such equipment.

In general terms, this disclosure is directed to an enhanced audio file generator. In some embodiments, an audio transformation model receives an input audio file representing speech and outputs an enhanced audio file with synthesized clean speech. In many embodiments, one or more machine learning models are used to transform an audio file to enhance the sound quality of the recording.

One aspect is a method of enhancing input speech in an input audio file, the method comprising receiving the input audio file representing the input speech, wherein the input audio file is recorded at an audio recording device and generating an enhanced audio file by applying an audio transformation model to the input audio file, wherein applying the audio transformation model to generate the enhanced audio file comprises extracting parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other preset parameters respectively defining other audio features, synthesizing clean speech based on the extracted parameters including the noise parameter, wherein synthesizing the clean speech comprises transforming the noise parameter to at least one defined value, and generating the enhanced audio file with the synthesized clean speech.

Another aspect is a method of enhancing input speech in an input audio file, the method comprising receiving the input audio file representing the input speech, wherein the input audio file is recorded at an audio recording device and generating an enhanced audio file by applying an audio transformation model to the input audio file, wherein applying the audio transformation model to generate the enhanced audio file comprises mapping the input audio file to a latent vector of audio features, wherein the audio transformation model comprises a transformation module that is trained to perform the mapping of the input audio file to the latent vector based on a decoder being enabled to synthesize clean speech from the latent vector, synthesizing the clean speech by applying the decoder to the latent vector, and generating the enhanced audio file with the synthesized clean speech.

Yet another aspect is an audio recording device comprising a processor in communication with a microphone, and a memory storing instructions, which when executed by the processor cause the audio recording device to record an input audio file to capture speech via the microphone, and generate an enhanced audio file by applying an audio transformation model to the input audio file, wherein to generate the enhanced audio file by applying the audio transformation model includes to extract parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other preset parameters respectively defining other audio features, synthesize clean speech based on the extracted parameters including the noise parameter, wherein to synthesize the clean speech comprises transforming the noise parameter to at least one defined value, and generate the enhanced audio file with the synthesized clean speech.

Another aspect is an audio recording device comprising a processor in communication with a microphone, and a memory storing instructions, which when executed by the processor cause the audio recording device to record an input audio file to capture speech via the microphone and generate an enhanced audio file by applying an audio transformation model to the input audio file, wherein to generate the enhanced audio file by applying the audio transformation model includes to map the input audio file to a latent vector of audio features, wherein the audio transformation model comprises a transformation module that is trained to map the input audio file to the latent vector based on a decoder being enabled to synthesize clean speech from the latent vector, synthesize the clean speech by applying the decoder to the latent vector, and generate the enhanced audio file with the synthesized clean speech.

Yet another aspect is a non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform receiving an input audio file representing input speech, wherein the input audio file is recorded at an audio recording device, and generating an enhanced audio file by applying an audio transformation model to the input audio file, wherein applying the audio transformation model to generate the enhanced audio file comprises extracting parameters defining audio features from the input audio file, the parameters including (i) a noise parameter defining noise in the input audio file and (ii) one or more other preset parameters respectively defining other audio features, synthesizing clean speech based on the extracted parameters including the noise parameter, wherein synthesizing the clean speech comprises transforming the noise parameter to at least one defined value, and generating the enhanced audio file with the synthesized clean speech.

Another aspect is a non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform receiving an input audio file representing the input speech, wherein the input audio file is recorded at an audio recording device and generating an enhanced audio file by applying an audio transformation model to the input audio file, wherein applying the audio transformation model to generate the enhanced audio file comprises mapping the input audio file to a latent vector of audio features, wherein the audio transformation model comprises a transformation module that is trained to perform the mapping of the input audio file to the latent vector based on a decoder being enabled to synthesize clean speech from the latent vector, synthesizing the clean speech by applying the decoder to the latent vector, and generating the enhanced audio file with the synthesized clean speech.

Various embodiments will be described in detail with reference to the drawings, wherein like reference numerals represent like parts and assemblies throughout the several views. Reference to various embodiments does not limit the scope of the claims attached hereto. Additionally, any examples set forth in this specification are not intended to be limiting and merely set forth some of the many possible embodiments for the appended claims.

In general terms, this disclosure is directed to an enhanced audio file generator. In many embodiments, a single machine learning model is used to transform an audio file to enhance the sound quality of the recording. In some embodiments, the method and systems disclosed herein allow users to record high quality voice recordings without using professional equipment. For example, a user can generate an original audio file on a mobile phone microphone or a connected Bluetooth headset. This original audio file is processed to extract features which are used to generate an enhanced audio file. In some embodiments, the enhanced audio file mimics features which are present in a professionally recorded and mixed audio file.

In some embodiments, the model for generating an enhanced audio file is computationally inexpensive and efficient allowing the model to run completely on the recording device (e.g., mobile device with a microphone) in real time. Additionally, the model for generating an enhanced audio file can be used in many different use cases such as: (1) a user recording a podcast, (2) a user recording an audio clip to interact with a podcast, artist, or other user, (3) a user recording an audio advertisement, and/or (4) speech recognition. Many other applications of the model for generating an enhanced audio file are discussed herein.

1 FIG. 100 102 102 104 104 106 104 108 104 102 110 110 112 114 116 112 116 118 illustrates an example environmentfor an enhanced audio file generator. In the embodiment shown, the enhanced audio file generatoris executed on the audio recording device. The audio recording devicerecords audio from the audio provider(e.g., recording the voice of a user of the audio recording device). An original audio filerecorded by the audio recording deviceis provided to the enhanced audio file generatorto generate an enhanced audio file. In some embodiments, the enhanced audio fileis uploaded via a network (e.g., the Internet) to a media delivery systemwhere it is stored in a media data storageas a media content item. The media delivery systemoperates to provide the media content itemamong other media content items to various consumer output devices, for example as part of a music streaming platform.

104 104 104 104 104 112 104 102 The audio recording deviceis a device with hardware and software components capable of recording audio. In some embodiments, the audio recording deviceis a single device (e.g., a smartphone). In some embodiments, the audio recording deviceis connected either wired or wirelessly to a device with a microphone. For example, a computing system connected to a microphone, or a smart phone connected to headphones having a microphone. In typical embodiments, the audio recording deviceincludes a processor which is in communication (integrated, wired, or wirelessly) with at least one microphone, and in electrical communication with a memory which stores instructions to perform various embodiments disclosed herein. Additionally, the audio recording deviceincludes a network interface and hardware to communicate with the media delivery system. In some embodiments, the memory stores instructions to cause the audio recording deviceto perform the applications of the enhanced audio file generatordescribed herein.

106 106 The audio provideris a user who generates the audio. In many of the embodiments described herein the audio providergenerates speech which is recorded. However, the systems and methods herein operate similarly for any type of audio. For example, the audio could be music from the audio provider's voice, an instrument, or a speaker. In other examples, the audio may be environmental noise (e.g., the sounds recorded at a park, beach, construction site, or any other location. Examples of types of content recorded include, podcasts, speech clips (e.g., speech clips responding/interacting with a podcast), and audio advertisements.

108 104 108 104 The original audio fileis an audio file which stores data recorded by the audio recording device. The systems and methods disclosed herein can be implemented using any of a variety of audio file formats. In some embodiments, the original audio fileis a waveform (WAV) audio file format file. In some embodiments, an audio recording device records audio from an audio provider in a WAV format. In other embodiments, the audio recording device first records the file in another format, such as MP3 and converts this file to the WAV audio file format. In some embodiments, different audio file formats can be used depending on the capabilities of the audio recording device.

102 108 108 102 102 102 2 5 FIGS.through 6 9 FIGS.through The enhanced audio file generatoroperates to process the original audio fileand generates or transforms the original audio fileinto the enhanced audio file. In some embodiments, the enhanced audio file generatoruses parameterized voice transformation as illustrated and described in reference to. In some embodiments, the enhanced audio file generatoruses latent voice transformation as illustrated and described in reference to. Additionally, the enhanced audio file generatorcan use any combination of the parameterized voice transformation and latent voice transformation.

110 102 108 110 110 108 110 108 108 110 The enhanced audio fileis generated by the enhanced audio file generator. In some embodiments, an audio transformation model is applied to the original audio fileto generate the enhanced audio file. In some embodiments, the enhanced audio fileis generated based on extracted features from the original audio file. In some of these embodiments, the enhanced audio fileis generated without directly referencing the original audio file. For example, the enhanced audio file may be generated based on features extracted from the original audio filewithout referencing the original audio file. In some embodiments, the enhanced audio filemimics features which are present in a professionally recorded and mixed audio file.

110 104 110 104 112 110 104 112 In some embodiments, the enhanced audio fileis temporarily stored on the audio recording device. In some embodiments, the enhanced audio fileis permanently stored on the audio recording device. In some examples, a user records audio which is not initially uploaded to the media delivery systembecause the user is not yet ready to share/publish yet the recorded audio. In some examples, a user may compile multiple enhanced audio files as part of the creation process for a media content item. In some embodiments, allowing a user to temporarily or permanently store the enhanced audio filereduces the amount of data transferred between the audio recording deviceand the media delivery system.

112 118 112 110 104 110 112 118 112 The media delivery systemoperates to provide media content to the consumer output devices. In the example shown, the media delivery systemfurther operates to receive an enhanced audio fileuploaded from the audio recording device. In some examples, the enhanced audio fileis a podcast which is uploaded to the media delivery systemso that it can be shared and played among the consumer output devices. In many embodiments, the media delivery systemincludes multiple servers which may be identical or similar and may provide similar functionality (e.g., to provide greater capacity and redundancy, or to provide services from multiple geographic locations). Alternatively, in these embodiments, some of the multiple servers may perform specialized functions to provide specialized services (e.g., services to enhance media content playback during travel, etc.). Various combinations thereof are possible as well.

114 110 The media data storagestores media content items. Examples of media content include audio content (e.g., songs, albums, podcasts, audio advertisements etc.). Many other examples of audio content are included within the scope of this disclosure including video content (e.g., the enhanced audio filecan be presented as the audio output for video content etc.).

116 110 116 116 116 106 116 118 118 The media content itemis, or includes at least a portion of, the enhanced audio fileuploaded to the media delivery system. In some embodiments, the media content itemcontent item is a podcast or segment to be inserted into a podcast. In other examples, the media content itemis an advertisement audio segment. In some embodiments, the audio provider can set the media content itemas private. In some embodiments, the audio providercan publish the media content item to the audio providers account to share the media content item. The consumer output devicestypically include one or more processors, a memory storing an application to perform various features including some of the features described herein and a speaker to output media content. The consumer output devicescan include a variety of I/O devices, computing devices, and software modules (including a software module for an operating system and software modules for interacting with and presenting media content).

118 116 112 118 104 118 104 The consumer output devicesare media playback devices which received media content items, including the media content itemfrom the media delivery system. Examples of consumer output devicesinclude smartphones, tablets, smart speakers, car audio systems, and other computing devices. In some embodiments, the audio recording deviceis one of the consumer output devices. For example, the user recording audio may also consume the media content item on the audio recording device.

104 112 2 4 FIGS.- In some embodiments, the audio recording deviceoperates with the media delivery systemfor training a model or models included in the enhanced audio file generator. In some embodiments, media content items with a high likelihood of clean speech segments are identified for training the model. For example, podcasts with lots of downloads, artwork, and many episodes are likely to have high quality clean speech segments. In some embodiments, these media content items are identified using a heuristic. The identified media content items are segmented into a plurality of segments and classified based on whether the segment includes music, speech, noise etc. The segments classified as containing clean speech are used as training examples for the audio transformation model. For example, using the techniques illustrated and described in. In some embodiments, the segments that are classified as noisy or as containing music are added back in as negative training examples.

2 FIG. 130 102 108 132 110 102 134 108 140 108 146 148 150 152 154 156 140 142 140 110 illustrates an example architecturefor an enhanced audio file generatorusing parameterized voice transformation. The enhanced audio file generator receives an original audio fileand (optionally) input recording environment dataand outputs an enhanced audio file. The enhanced audio file generatorincludes a speech analyzerwhich decodes the original audio fileto extract the parameters. The parameters can include various audio features extracted from the original audio fileincluding any combination of phonemes, pitch salience, voice timbre, noise, voice volume, and latent features. The parametersare provided to the speech synthesizerwhich processes the parametersto generate the enhanced audio file.

108 110 1 FIG. Examples of the original audio fileand the enhanced audio fileare illustrated and described in reference to.

132 108 132 108 108 132 108 108 132 132 132 The input recording environment dataincludes data which supplements the original audio file. For example, the input recording environment datacan include information related to how the original audio filewas recorded, such as the type of device the original audio filewas recorded on, a type of microphone, a connection type of the microphone (e.g., integrated, wired, or wireless), etc. In some embodiments, the input recording environment dataincludes data such as user account data associated with the original audio file, a location where the original audio file was recorded, whether the original audio filewas recorded in a professional studio, metadata associated with the original audio file, etc. In some embodiments, the input recording environment datais automatically generated. For example, the audio recording device may include an application which determines system information of the audio recording device, information of connected devices, user account information, device location information, or any combination thereof to automatically generate the input recording environment data. In other embodiments, some or all of the input recording environment datais manually provided by a user.

134 140 108 140 134 140 108 134 140 134 The speech analyzerextracts parametersfrom the original audio file. In some embodiments, the parametersare preset parameters that correspond to audio features. In some embodiments, the speech analyzeruses a neural network to extract some or all of the parametersfrom the original audio file. In some embodiments, the speech analyzeris trained to extract the parameters. In some embodiments, the speech analyzeris trained to extract the parameters which produce individual components, one for each feature, and summing the components.

140 134 140 152 110 In some embodiments, the parametersare preset parameters that correspond to audio features which the speech analyzeris trained to identify and calculate. In some embodiments, the parameters are transformed to provide a desired effect. For example, the parameterscan be transformed to match or mimic features in professionally recorded and mixed audio. For example, the noisecan be transformed to zero to generate clean speech in the enhanced audio file.

140 146 148 150 152 154 156 146 148 148 150 152 154 156 140 4 5 FIGS.and In some embodiments, the parametersinclude any combination of phonemes, pitch salience, voice timbre, noise, voice volume, and latent features. Phonemesincludes perceptually distinct units of sound. Pitch Salienceincludes a measure of tone sensation. In some embodiments, pitch salienceincludes a measure of the predominance of different frequencies in an audio single at every time frame. Voice Timbreincludes a measure of the global sound quality. In some embodiments, noiseincludes the noise identified in the original audio file. Examples of noise includes background noise. Voice volume includesincludes a measure of voice volume at the different time periods in the original audio file. Latent featuresincludes any other features extracted from the speech analyzer. For example, latent features may extract breathing sounds from the original audio file. In some embodiments, the latent features are encoded in a latent vector with a transformation module and decoder trained to map features to the latent vector according to the example embodiment illustrated in. In some embodiments, the parameterscan be adjusted based on other models to create a desired effect. For example, the pitch salience can be adjusted by the output of another model to create an input speech to output singing effect.

142 110 140 142 110 108 142 110 140 142 110 140 132 142 3 FIG. The speech synthesizergenerates the enhanced audio filebased on the parameters. The speech synthesizerreconstructs the enhanced audio filewithout reference to the original audio file. For example, the speech synthesizercan generate the enhanced audio filebased only on the parameters. In the embodiment shown, the speech synthesizergenerates the enhanced audio filebased on only the parametersand the input recording environment data. An example of the architecture for the speech synthesizeris illustrated and described in.

3 FIG. 2 FIG. 142 142 142 illustrates an example architecture for a speech synthesizer. The speech synthesizeris an example of the speech synthesizerillustrated and described in reference to.

142 146 148 150 152 154 156 132 142 186 110 152 152 2 FIG. 1 2 FIGS.and Inputs to the speech synthesizerinclude phonemes, pitch salience, voice timbre, noise, voice volume, latent features, and input recording environment data. Details for these inputs are illustrated and described in reference to. The speech synthesizeroutputs a reconstructed audio file. In some embodiments, to generate an enhanced audio file (e.g., the enhanced audio fileillustrated and described in) the noiseparameter is set to at least one defined value (e.g., zero, a non-zero constant, or a value that may vary over time), which is inaudible or otherwise deemed acceptable for clean speech, such that the reconstructed audio file includes audio data representing clean speech. For example, the noiseparameter can be set to a value which may vary over time but remains inaudible to a user or is otherwise deemed an acceptable level of noise for clean speech.

132 150 146 148 152 154 156 156 In typical embodiments, the input recording environment dataand voice timbreare global inputs (e.g., inputs which do not vary over time) while phonemes, pitch salience, noise, voice volume, and latent featuresvary over time. In some embodiments, latent featuresinclude global features as well as, or instead of, time varying features.

142 180 180 182 146 148 150 156 132 182 180 3 FIG. The speech synthesizerincludes neural network blocks. The neural network blocksuse the received inputs as basis to generate the synthesized speech, which may be unleveled as shown in(e.g., the unleveled synthesized speech), based on the received inputs. In some embodiments, the neural network blocks receive any combination of parameters, such as the phonemes, pitch salience, voice timbre, and latent featuresas well as the input recording environment data. which generates unleveled synthesized speech. In some embodiments, the neural network blocksare trained using a supervised machine learning technique.

154 152 180 180 182 182 154 184 152 184 186 182 152 184 186 152 142 186 186 180 142 180 In some embodiments, the voice volumeand noiseare used as inputs during the training of the neural network blocks. In some examples, unleveled speech and/or noise are features which a user would like to remove from a recording. The neural network blocksare trained to generate the unleveled synthesized speech. In some embodiments, the unleveled synthesized speechis point-wise multiplied by the extracted voice volumeto generate the synthesized speech. In some embodiments, noiseis added to the (e.g., leveled) synthesized speechto generate the reconstructed audio file. In some embodiments, the leveling with the voice volume is option and the noise is added to the synthesized speech (e.g., the unleveled synthesized speech) generated by the neural network. In some embodiments, adding the noiseto the synthesized speechis done to train the neural network to accurately reconstruct the original audio file without noise. In these embodiments, the noise is added back during the training stage so the reconstructed audio filematches the input file. In some embodiments, the noiseis ultimately set to zero once the training is complete and the speech synthesizeris being used to generate clean speech. The reconstructed audio fileis then compared to the original audio file. This process repeats until the differences between the reconstructed audio fileand the original audio file are below a threshold. At this point, the neural network blocksare trained to generate synthesized speech which is leveled and without noise. In this manner, the speech synthesizerforces the neural network blocksto learn how to generate leveled speech without noise.

180 142 180 186 142 180 154 152 In some embodiments, the architecture operates with two paths one path being used to train the neural network blocks(e.g., as described above) and a second path which is used when applying the trained speech synthesizer. For example, the trained neural network blocksmay directly provide the output reconstructed audio filewith clean speech when the trained speech synthesizeris being applied in an application. However, the architecture shown functions as a single path. For example, once the neural network blocksare trained the voice volumeinput is set to a constant (e.g., a vector of ones) and the noiseis set to a constant (e.g., 0). In some embodiments, the noise is set to a non-zero level which is inaudible to a user or otherwise deemed an acceptable level of noise for clean speech. This results in an audio file with enhanced speech being generated as the output. In some embodiments, the noise is assigned a value which may vary over time but remains inaudible to a user or is otherwise deemed an acceptable level of noise for clean speech.

180 180 150 180 186 146 156 132 132 180 180 In addition to, or instead of, training the neural network blocks to generate level speech and remove background the other parameters can be used to train the neural network blocks. For example, in some use cases transforming the voice timbre may be desired. In these embodiments, the neural network blocksare trained in a similar manner to the voice volume and background noise. For example, the neural network blockscan be trained to transform voice timbreby receiving three examples of extracted voice timbre from audio samples. Two of the samples may be from a voice with desired voice timbre and the third with a different voice timbre. The neural network blocksare trained until the reconstructed audio fileoutputs a voice timbre closer to the two examples with the desired voice timbre. In another example, pitch salience can be extracted and reintroduced to the synthesized speech to train the neural network blocks to remove certain pitch salience features. This technique can be repeated for phonemes(e.g., to remove hard “s” or “p” sounds), latent features, and input recording environment data. In some embodiments, the input recording environment datacan be used to train the neural network blocksto remove common abnormalities captured on specific recording devices (e.g., a certain type of microphone may struggle to capture certain frequencies which the neural network blockscan be trained to removed).

108 186 186 132 In some embodiments, a single model is designed to use a loss function which is the sum of several components. One component is audio reconstruction which is calculated as the mean squared error between the original audio fileand the reconstructed audio file. A second component is a phoneme component which is calculated by the mean squared error between estimated phonemes and the output of an already trained phoneme estimation model. A third component is a pitch salience component which calculates the mean squared error between the estimated pitch salience and an already trained pitch salience model. A fourth component is a voice timbre component which uses a triplet loss technique which analyzes multiple target samples (typically two) and a negative sample (typically one) and determines whether the reconstructed audio fileis closer to the target sample or the negative sample. A fifth component is a noise estimation (e.g., a background noise estimation) which is the mean squared error between the estimated noise and the actual noise. These components are combined to generate a single transformation model. In some embodiments, the date required to train components includes the original audio file and input recording environment data. These inputs are used over several training traces to generate the following outputs: (1) synthesized speech; (2) noise; (3) a pitch salience neural network on the synthesized clean speech; (4) output from a phoneme estimation model on the synthesized speech; (5) a second synthesized speech sample from same recording; and (6) a sample from a different recording.

4 FIG. 187 187 188 189 190 191 illustrates an example methodof enhancing input speech in an input audio file. The methodincludes the operations,,, and.

188 187 The operationreceives an input audio file representing input speech. In some embodiments, the input audio file is recorded at an audio recording device. In some embodiments, the methodreceives and processes the input audio file in real-time as the user is recording speech at the audio recording device.

189 190 191 192 5 FIG. In some embodiments, the operations,, andare part of a step for generating an enhanced audio file by applying an audio transformation model to the input audio file. In some embodiments, the audio transformation model is trained using the methodillustrated and described in reference to.

189 189 The operationextracts parameters defining audio features from the input audio file. In some embodiments, the parameters include a noise parameter defining noise in the input audio file and one or more other preset parameters respectively defining other audio features. In some embodiments, the one or more other preset parameters respectively define one or more of phonemes, pitch salience, voice timbre, voice volume, or any combination thereof. In some embodiments, the operationfurther determines input recording environment data from the audio recording device.

190 The operationsynthesizes clean speech based on the extracted parameters. In some embodiments, the extracted parameters include a noise parameter and synthesizing the clean speech includes transforming the noise parameter to at least one defined value. In some embodiments, the at least one defined value is set to zero causing the noise to inaudible. Alternatively, the noise parameter can be set to a non-zero level which is inaudible to a user or otherwise deemed an acceptable level for clean speech. In some embodiments, the clean speech is synthesized using a neural network. In some embodiments, the clean speech is synthesized without referencing the input audio file. In some embodiments, the audio transformation model accounts for input recording environment data of the audio recording device.

191 110 1 2 FIGS.and The operationgenerates the enhanced audio file with the synthesized clean speech. In some embodiments, the enhanced audio file is generated without referencing the input audio file. In some embodiments, the enhanced audio file is the enhanced audio fileillustrated and describe in reference to.

195 195 195 187 In some embodiments, an audio recording device comprising, a processor in communication with a microphone and a memory storing instructions, which when executed by the processor cause the audio recording device to perform the method. In some embodiments, the methodis performed entirely on the audio recording device. In some embodiments, the methodis performed in real-time as a user records audio at the audio recording device. In some embodiments, A non-transitory computer-readable storing instructions which, when executed by one or more processors, cause the one or more processors to perform the method.

5 FIG. 2 3 FIGS.and 192 192 142 192 193 194 195 196 197 illustrates an example methodfor training the audio transformation model. In some embodiments, the methodis used to train a neural network to synthesize clean speech based on preset parameters for audio features as part of the speech synthesizeras shown in. The methodincludes the operations,,,, and.

193 The operationreceives a training audio file. The training audio file includes audio data representing training speech. In some embodiments, the training audio file includes noise and a corresponding target audio file is not required or used in the training of the audio transformation model.

194 The operationextracts training parameters from the training audio file. In some embodiments, the training parameters include a noise parameter defining noise in the training audio file. In some embodiments, the noise includes background noise detected in the audio file.

195 The operationsynthesizes reconstructed speech. In some embodiments the reconstructed speech is synthesized using a neural network. In some embodiments, the different preset parameters correspond to several components which the neural network receives as inputs to synthesizes the reconstructed speech.

196 The operationgenerates an output audio file with the reconstructed speech. In some embodiments, the noise parameter is added included in the output audio file with the reconstructed speech. This trains the neural network to reconstruct the speech without the noise, such that when the noise is added back into the output audio file the output audio file will match the training audio file when the neural network reconstructs speech without noise.

197 The operationcompares the output audio file generated with the reconstructed speech to the training audio file including the training speech. This comparison is used to train and further refine the neural network to synthesize clean speech.

2 5 FIGS.through Advantages of the parameterized voice transformation architecture illustrated and described in reference toinclude: (1) generating a single model which transforms an original audio file to an enhanced audio file based on any of a variety of parameters, (2) the single model is computationally inexpensive so a user can download the model at a mobile computing device allowing a user to generate the enhanced audio file without uploading audio to a server, (3) the single model can process audio in real time on many different features with one click; (4) the model can be generated on any corpus of training data (no need to for paired clean/noisy speech training samples). Because, in some embodiments, the single model can be downloaded and run on a user device the enhanced audio can be generated anywhere (e.g., at locations with no network connectivity) and/or without the privacy concerns of uploading speech to a server.

6 FIG. 7 FIG. 200 202 illustrates an example architecturefor an initial stage of training a latent vector transformation modelto generate an enhanced audio file. In some embodiments, a subsequent stage (illustrated and described in reference to) is performed after completing the initial stage.

202 204 206 207 207 210 202 208 206 207 208 207 202 204 210 202 Initially the latent vector transformation modelis trained on clean speech. The encodermaps the clean audio to a latent vectorof audio features. The decoder uses the latent vectorto reconstruct the audio and output clean speech. By encoding and decoding features the latent vector transformation modellearns which speech features are relevant (e.g., features such as pitch, volume, timbre, etc.) and the decoderlearns how to generate speech when given these audio features. The encoderlearns which weights to apply to which features to encode the latent vectorand the decoderlearns which weights to use to decode the latent vector. The latent vector transformation modelis able to freely identify audio features in the latent space. After the model learns to encode and decode these features such that the differences between the clean speechand the output clean speechis below a threshold the initial stage of training the latent vector transformation modelis complete.

7 FIG. 6 FIG. 218 202 illustrates an example architecturefor a subsequent stage of training a latent vector transformation modelto generate an enhanced audio file. The subsequent stage is performed after the completion of the initial stage illustrated and described in reference to.

206 208 202 220 230 220 224 224 222 224 222 222 224 224 220 207 208 228 228 230 6 FIG. 6 FIG. At the subsequent stage, the transformer is initialized with the weights for encoding audio features in a latent vector from the encoderas described in. The weights for the decoderas trained inare frozen. In some embodiments, the latent vector transformation modelis further trained at this stage using a noisy speech training exampleand a paired target clean speech training example. The noisy speech training exampleis provided to the transformation module. In some embodiments, the transformation modulereceives a conditioninput to condition the transformed based on known features. For example, the transformation modulemay receive a conditioninput which indicates that recording was performed on a specific device (e.g., a type of phone) or using certain hardware (e.g., a Bluetooth microphone). In some embodiments, the conditioninput indicates a noise type present in the input audio file. For example, data associated with noisy speech may be used to condition the transformation moduleto learn how to decode audio features with different types of noise. The transformation modulelearns to map the noisy speech training exampleinto the latent vectorin a way that the decodercan reconstruct clean speech which is output as output clean speech. The output clean speechis compared with the target clean speech training exampleto supervise the training of the transformation module.

202 220 230 224 208 208 202 2 5 FIGS.through Advantages of training the latent vector transformation modelover these two stages include lowering the number of paired samples of noisy speech training exampleand target clean speech training examplerequired to train the transformation moduleas the decoderis trained at a separate stage which can use unpaired samples of clean speech. Additionally, the decoder learns to reconstruct output on a lot of clean samples improving the performance of the decoder. In some embodiments, the latent vector transformation modelis used to map latent features in reference to some of the embodiments illustrated and described in.

224 In some embodiments, the paired noisy speech and clean training examples used for training the transformation modulecan be artificially generated by reducing the quality of clean speech training examples. For example, a clean speech training example can be down sampled, encoded to a lossy audio format (decoding to a lossy format such as MP3), randomly adjusting audio overtime, identifying time periods in the sample with “ess” sounds and/or “p” sounds and adjusting the volume at these time periods, overdriving the signal, applying equalization curves to simulate bad microphones/recording conditions, adding reverberation, mixing with different kinds of background noises (mouth noise, street hum, wind, etc.).

8 FIG. 240 240 242 244 246 248 illustrates an example methodfor enhancing input speech in an input audio file. The methodincludes the operations,,, and.

242 The operationreceives an input audio file. In some embodiments, the input audio file includes audio data representing input speech. In some embodiments, the input speech is recorded at an audio recording device.

244 246 248 260 9 FIG. In some embodiments, the operations,, andare part of a step for generating an enhanced audio file by applying an audio transformation model to the input audio file. In some embodiments, the audio transformation model is trained using the methodillustrated and described in reference to.

244 The operationmaps the input audio file to a latent vector of audio features with a transformation module. In some embodiments, the audio transformation model comprises a transformation module that is trained to perform the mapping of the input audio file to the latent vector based on a decoder being enable to synthesize clean speech from the latent vector.

246 262 260 9 FIG. The operationsynthesizes clean speech by applying the decoder to the latent vector. The decoder decodes the latent vector to audio data with speech. In some embodiments, the decoder is trained to synthesize clean speech by training an encoder to map clean speech training examples on the latent vector and training the decoder to reconstruct the clean speech training examples (e.g., the operationof the example methodillustrated and describe in reference to). In some embodiments, the clean speech is synthesized without referencing the input audio file. In some embodiments, the audio transformation model accounts for input recording environment data of the audio recording device.

248 110 1 FIG. The operationgenerates the enhanced audio file with the synthesized clean speech. In some embodiments, the enhanced audio file is generated without referencing the input audio file. In some embodiments, the enhanced audio file is the enhanced audio fileillustrated and described in reference to.

240 240 240 240 In some embodiments, an audio recording device comprising, a processor in communication with a microphone and a memory storing instructions, which when executed by the processor cause the audio recording device to perform the method. In some embodiments, methodis performed entirely on the audio recording device. In some embodiments, the methodis performed in real-time a s a user records audio at the audio recording device. In some embodiments, A non-transitory computer-readable medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform the method.

9 FIG. 8 FIG. 260 240 260 262 264 266 illustrates an example methodfor training a latent vector transformation module. In some embodiments, the latent vector transformation module is the audio transformation model applied in the methodas shown in. The methodincludes the operations,, and.

262 264 262 264 266 6 FIG. The operationtrains an encoder to map clean speech training examples to the latent vector and the operationtrains a decoder to reconstruct the clean speech training examples. In some embodiments, the operationsandare sequentially performed multiple times in order to train the decoder to reconstruct speech. For example, as illustrated and described in the example architecture illustrated and described in reference to. In some embodiments, the decoder is trained prior to the training of the transformation module and the decoder is used for training the transformation module at the operation.

266 266 7 FIG. The operationtrains a transformation module to map noisy speech training examples on the latent vector. An example architecture for the operationis illustrated and described in reference to. In some embodiments, training the transformation module comprises training the transformation module to map noisy speech training examples on the latent vector such that the decoder reconstructs clean speech output matching clean speech training examples paired with the noisy speech training examples. In some embodiments, the noisy speech training examples are artificially produced from the paired clean speech training examples. In some embodiments, the clean speech training examples include speech which was professionally recorded and mixed.

10 FIG. Example 1 illustrates a user generating a studio-quality podcast. The user can record the podcast anywhere, including noisy environments and process the recording with the enhanced audio generator to generate a studio-quality recording which can be uploaded and shared with listeners of the podcast. Example 2 illustrates an example of uploading a short clip to a social media platform. In some embodiments, a podcaster may want to interact with listeners. In this example, the users can upload enhanced audio files which reduces the amount of time required to integrate the user submitted content into podcasts. Additionally, the podcaster can use more user submitted content because the content will be of a sufficiently high quality. In the example shown, the noisy short voice recording is provided to the enhanced audio generator which generates a de-noised recording which the user can upload to a social media platform. Example 3 illustrates a use case with a speech recognition system (e.g., a voice command device or a voice assistant). For example, a user may provide a voice command with a lot of noise (e.g., due to environmental sounds or a microphone issue) and the enhanced audio generator and generate an enhanced speech recording which allows the voice assistant to improve the accuracy and processing speed for processing the voice command at a speech recognition system. Example 4 illustrates an example for recording and generating a studio quality recording for a song. In some embodiments, an enhanced audio generator is integrated in a mixing application which allows a user to mix studio quality voice recordings with music. Example 5 illustrates an example for producing a professional sounding advertisement. The user provides a voice over recording for an advertisement and the enhanced audio generator generates a studio-quality recording. In some embodiments, the enhanced audio generator is built into a mixing application which allows a user to add background music to quickly create a professional sounding advertisement. illustrates example applications for the enhanced audio file generator. Many other examples are described herein or are included within the scope of the disclosure.

Further example applications of the enhanced audio file generator include: (1) speech denoising (e.g., replace background noise component with silence before synthesis) (2) studio quality recording generation; (3) voice beautification (e.g., replace voice timbre and pitch salience components with ones transformed by a different neural network); (4) voice swapping (e.g., replace the voice timbre component with the voice timbre component produced by the speech analyzer for a different voice recording); (5) accent swapping (e.g., replace the pitch salience and phoneme components with ones transformed by a different neural network); (6) voice anonymization (replace the voice timbre component with the voice timbre component produced by the speech analyzer for a different voice recording, e.g., by swapping these features with a voice that does not sound like the user or a random voice); (7) explicit word scrubber (e.g., by identifying explicit words/sounds using a separate model, and remove-set to zero- or replace these words/sounds, before synthesis); (8) word removal or sound removal (e.g., by identifying words/sounds and removing these sounds as part of the enhanced file generation); (9) speech to singing or singing to speech application (e.g., replace the pitch salience component with one generated by another model); (10) age morphing (e.g., replace the voice timbre and pitch salience components with one generated by another model; and (11) nonhuman voice morphing (e.g., replace the voice timbre, pitch salience, and phoneme components with ones generated by another model).

11 FIG. 302 304 306 302 308 308 illustrates example user interfaces for an application which uses the enhanced audio file generator. The user interfaceis presented as part of an application which lets users record audio. The user interfaceis presented to a user while the recording is active. In some embodiments, the enhanced audio generator automatically generates enhanced audio as the user records. Because the enhanced audio generator can process audio faster than the audio is recorded, in some embodiments, the enhanced audio file is generated in real time. However, in the embodiment shown at the user interfacethe user selects a setting for enhanced audio after the recording is complete. The user can set the audio enhancement setting with a single selection. In alternative embodiments, the enhanced audio selection is presented on the user interfaceto process the original audio file in real time. The user interfaceis presented to a user after the audio enhancement is complete. The user interfaceincludes a selection to add background music. This can allow the user to quickly generate professional sounding audio (e.g., songs, podcasts, audio advertisements, etc.).

While various example embodiments of the present invention have been described above, it should be understood that they have been presented by way of example, and not limitation. It will be apparent to persons skilled in the relevant art(s) that various changes in form and detail can be made therein. Thus, the present invention should not be limited by any of the above described example embodiments, but should be defined only in accordance with the following claims and their equivalents.

The example embodiments described herein may be implemented using hardware, software or a combination thereof and may be implemented in one or more computer systems or other processing systems. However, the manipulations performed by these example embodiments were often referred to in terms, such as entering, which are commonly associated with mental operations performed by a human operator. No such capability of a human operator is necessary, in any of the operations described herein. Rather, the operations may be completely implemented with machine operations. Useful machines for performing the operation of the example embodiments presented herein include general purpose digital computers or similar devices.

From a hardware standpoint, a CPU typically includes one or more components, such as one or more microprocessors, for performing the arithmetic and/or logical operations required for program execution, and storage media, such as one or more disk drives or memory cards (e.g., flash memory) for program and data storage, and a random access memory, for temporary data and program instruction storage. From a software standpoint, a CPU typically includes software resident on a storage media (e.g., a disk drive or memory card), which, when executed, directs the CPU in performing transmission and reception functions. The CPU software may run on an operating system stored on the storage media, such as, for example, UNIX or Windows (e.g., NT, XP, Vista), Linux, and the like, and can adhere to various protocols such as the Ethernet, ATM, TCP/IP protocols and/or other connection or connectionless protocols. As is well known in the art, CPUs can run different operating systems, and can contain different types of software, each type devoted to a different function, such as handling and managing data/information from a particular source, or transforming data/information from one format into another format. It should thus be clear that the embodiments described herein are not to be construed as being limited for use with any particular type of server computer, and that any other suitable type of device for facilitating the exchange and storage of information may be employed instead.

A CPU may be a single CPU, or may include multiple separate CPUs, wherein each is dedicated to a separate application, such as, for example, a data application, a voice application, and a video application. Software embodiments of the example embodiments presented herein may be provided as a computer program product, or software, that may include an article of manufacture on a machine accessible or non-transitory computer-readable medium (i.e., also referred to as “machine readable medium”) having instructions. The instructions on the machine accessible or machine readable medium may be used to program a computer system or other electronic device. The machine-readable medium may include, but is not limited to, floppy diskettes, optical disks, CD-ROMs, and magneto-optical disks or other type of media/machine-readable medium suitable for storing or transmitting electronic instructions. The techniques described herein are not limited to any particular software configuration. They may find applicability in any computing or processing environment. The terms “machine accessible medium”, “machine readable medium” and “computer-readable medium” used herein shall include any non-transitory medium that is capable of storing, encoding, or transmitting a sequence of instructions for execution by the machine (e.g., a CPU or other type of processing device) and that cause the machine to perform any one of the methods described herein. Furthermore, it is common in the art to speak of software, in one form or another (e.g., program, procedure, process, service, application, module, unit, logic, and so on) as taking an action or causing a result. Such expressions are merely a shorthand way of stating that the execution of the software by a processing system causes the processor to perform an action to produce a result.

While various example embodiments have been described above, it should be understood that they have been presented by way of example, and not limitation. It will be apparent to persons skilled in the relevant art(s) that various changes in form and detail can be made therein. Thus, the present invention should not be limited by any of the above described example embodiments, but should be defined only in accordance with the following claims and their equivalents.

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

February 24, 2026

Publication Date

July 2, 2026

Inventors

Rachel Malia Bittner
Jan Van Balen
Daniel Stoller
Juan José Bosch Vicente

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “Enhanced Audio File Generator” (US-20260188337-A1). https://patentable.app/patents/US-20260188337-A1

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.