The disclosed technology includes a voicemail translation service of a telecommunications network. The voicemail translation service can receive a voicemail message communicated from a calling party to a called party. The voicemail translation service calls an application programming interface (API) to upload the voicemail message to a transcription service, which completes transcription of the voicemail message in a default language. In response to determining that the default language of the completed transcription does not match a target language of the called party, the voicemail translation service triggers another API to generate a translation of the transcription in accordance with the target language. The voicemail translation service then stores the translated transcription in a voicemail storage system that is accessible by the called party.
Legal claims defining the scope of protection, as filed with the USPTO.
receiving, via the telecommunications network, a voicemail message communicated from a first user device associated with a first user to a second user device associated with a second user; uploading the voicemail message to a transcription service; in response to the voicemail message being uploaded to the transcription service, initiating a transcription of the voicemail message by the transcription service; receiving, from the transcription service, a completed transcription of the voicemail message in a default language value; wherein the target language value corresponds to a user-chosen language value in a user profile of the second user; retrieving, from a database, an indication of a target language value for the completed transcription, comparing the default language value of the completed transcription to the target language value; and in response to determining that the default language value of the completed transcription does not match the target language value, generating a translation of the completed transcription in accordance with the target language value. . A method performed by a voicemail translation service of a telecommunications network configured to translate a voicemail message, the method comprising:
claim 1 wherein a language value of the translation of the completed transcription matches the target language value. storing the translation of the completed transcription in a voicemail storage system accessible by the second user, . The method offurther comprising:
claim 1 wherein the confirmation signal indicates a successful receipt and processing of an entirety of the voicemail message. receiving a confirmation signal from the transcription service, . The method offurther comprising, prior to initiating the transcription of the voicemail message by the transcription service:
claim 1 inspecting progress of the transcription of the voicemail message by the transcription service; receiving an indication of an incomplete status for the transcription of the voicemail message by the transcription service; and in response to receiving the indication of the incomplete status, inspecting the progress of the transcription of the voicemail message after a time interval and until the transcription of the voicemail message is completed. . The method offurther comprising:
claim 1 wherein the user preference corresponds to a user-selected delivery method for notifications of translated voicemail messages. retrieving, from the database, an indication of a user preference, . The method offurther comprising:
claim 5 wherein a language value of the translation of the completed transcription matches the target language value; storing the translation of the completed transcription in a voicemail storage system accessible by the second user, determining that the user preference includes a short message service (SMS); and transmitting, to the second user device, an SMS message including a notification that the translation of the completed transcription is stored at the voicemail storage system. . The method offurther comprising:
claim 5 wherein a language value of the translation of the completed transcription matches the target language value; storing the translation of the completed transcription in a voicemail storage system accessible by the second user, determining that the user preference includes an electronic mailbox; and wherein the secure push proxy server delivers the push notification to the electronic mailbox of the second user, and wherein the push notification indicates that the translation of the completed transcription is stored at the voicemail storage system. transmitting a push notification to a secure push proxy server, . The method offurther comprising:
claim 5 wherein a language value of the translation of the completed transcription matches the target language value; storing the translation of the completed transcription in a voicemail storage system accessible by the second user, determining that the user preference includes electronic mail; and wherein the message transfer agent delivers the notification to the second user device via a message delivery agent. transmitting a notification about the translation of the completed transcription is stored at the voicemail storage system to a message transfer agent, . The method offurther comprising:
claim 2 generating a notification indicating that the translation of the completed transcription is stored at the voicemail storage system; and communicating, via the telecommunications network, the notification to the second user device. . The method offurther comprising:
claim 1 removing the uploaded message and completed transcription from the transcription service. . The method offurther comprising:
at least one hardware processor; and receive, via a telecommunications network, a voicemail message for a user of a user device; cause a transcription service to transcribe the voicemail message in a default language value; retrieve an indication of a target language value for the user of the user device; in response to determining that the default language value of the transcription does not match the target language value, generate a translation of the transcription in accordance with a target language value; inspect progress of the transcription of the voicemail message by the transcription service; receive an indication of an incomplete transcription of the voicemail message by the transcription service; and in response to receiving the indication of the incomplete transcription, inspect the progress of the transcription of the voicemail message after a time interval and until the transcription of the voicemail message is completed. at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to: . A system comprising:
claim 11 wherein a language value of the translation of the transcription matches the target language value. store the translation of the transcription in a voicemail storage system accessible by the user, . The system offurther caused to:
claim 11 wherein the user preference corresponds to a user selected delivery method for notifications of translated voicemail messages. retrieve, from a database of user profiles, an indication of a user preference, . The system of, further caused to:
claim 13 wherein a language value of the translation of the transcription matches the target language value; store the translation of the transcription in a voicemail storage system accessible by the user, determine that the user preference includes a short message service (SMS); and transmit, to the user device, an SMS message including a notification about the translation of the transcription being stored at the voicemail storage system. . The system of, further caused to:
claim 13 wherein a language value of the translation of the transcription matches the target language value; store the translation of the transcription in a voicemail storage system accessible by the user, determine that the user preference includes an electronic mailbox; and wherein the secure push proxy server delivers the push notification to the electronic mailbox of the user, and wherein the push notification indicates that the translation of the transcription is stored at the voicemail storage system. transmit a push notification to a secure push proxy server, . The system of, further caused to:
claim 13 wherein a language value of the translation of the transcription matches the target language value; store the translation of the transcription in a voicemail storage system accessible by the user, determine that the user preference includes electronic mail; and transmit a notification indicating that the translation of the transcription is stored at the voicemail storage system to a message transfer agent. . The system of, further caused to:
claim 12 wherein a language value of the translation of the transcription matches the target language value; store the translation of the transcription in a voicemail storage system accessible by the user, generate a notification indicating that the translation of the transcription is stored at the voicemail storage system; and communicate, via the telecommunications network, the notification to the user device. . The system of, further caused to:
receive an indication of a voicemail message for a user of a user device; cause a transcription service to transcribe the voicemail message in a default language; retrieve an indication of a target language for the user; compare the default language of the transcription to the target language; determining that the default language does not match the target language; generate a translation of the transcription in accordance with the target language; inspect progress of the transcription of the voicemail message by the transcription service; receive an indication of an incomplete status for the transcription of the voicemail message by the transcription service; and in response to receiving the indication of the incomplete status, inspect progress of the transcription of the voicemail message; and store the translation of the transcription in a voicemail storage system accessible by the user. . A non-transitory, computer-readable storage medium comprising instructions recorded thereon, wherein the instructions, when executed by at least one data processor of a system, cause the system to:
claim 18 wherein the user preference indicates to a user-selected notification method for translated voicemail messages. retrieve, a user profile of the user, an indication of a user preference, . The non-transitory, computer-readable storage medium of, further causing the system to:
claim 18 communicate, to the user device, a notification indicating that the translation of the transcription is stored at the voicemail storage system. . The non-transitory, computer-readable storage medium of, further causing the system to:
Complete technical specification and implementation details from the patent document.
This application is a continuation of U.S. patent application Ser. No. 18/492,598, filed Oct. 23, 2023, which is hereby incorporated by reference in its entirety.
Machine translation is use of either rule-based or probabilistic (e.g., statistical and, most recently, neural network-based) machine learning approaches to translation of text or speech from one language to another, including the contextual, idiomatic and pragmatic nuances of both languages.
Unedited machine translation is publicly available through tools on the Internet such as Google Translate, Babylon, DeepL Translator, and StarDict. These tools produce rough translations that, under favorable circumstances, “give the gist” of the source text. With the Internet, translation software can help non-native-speaking individuals understand webpages published in other languages. Whole-page-translation tools are of limited utility, however, since they offer only a limited potential understanding of the original author's intent and context; translated pages tend to be more erroneously humorous and confusing than enlightening.
The technologies described herein will become more apparent to those skilled in the art from studying the Detailed Description in conjunction with the drawings. Embodiments or implementations describing aspects of the invention are illustrated by way of example, and the same references can indicate similar elements. While the drawings depict various implementations for the purpose of illustration, those skilled in the art will recognize that alternative implementations can be employed without departing from the principles of the present technologies. Accordingly, while specific implementations are shown in the drawings, the technology is amenable to various modifications.
The disclosed technology includes a system that provides users with the ability to receive voicemail transcriptions in a user-specific language. The system can operate with telecommunications services including communications between at least a pair of user devices (e.g., smartphones). The system includes a server that manages the process of delivering translated voicemail transcriptions to a user device. For example, the server can receive the voicemail from a first smartphone and create a transcription of the voicemail by calling an Application Programming Interface (API) to do so. The server can generate a translation of the transcription and provide the translated transcription to a second smartphone.
The system can provide users with automatic translated transcriptions by utilizing natural language processing (NLP) and machine learning to transcribe voicemails. In one example, the system utilizes automatic speech recognition (ASR) technology to transcribe the spoken words from the voicemail into written text. A model of the system is able to determine the language spoken in the voicemail by recognizing language-specific acoustic features after being trained on large amounts of speech data from a wide array of different languages. After transcribing the voicemail, NLP techniques are applied to preprocess the text by correctly identifying words in cases of homophones and correctly identifying which punctuation to use in the transcription of the voicemail. After finalizing the transcription, the server can compare the detected language with a user-selected language. In response to determining the detected language does not match the user-selected language, the system can call the API to translate the transcription using translation models. In one example, the system utilizes neural networks trained for language translation such as a neural machine translation algorithm to translate the transcription. The system can account for language-specific rules such as grammar rules and common expressions; therefore, increasing the accuracy of the generated translations.
By calling an API to translate a transcription of a received voicemail, the system simplifies the process of translating the transcription. In addition to that, the API allows the system to reduce development time by interacting with pre-existing software and allowing the system to easily adapt to utilizing those software programs. Overall, the API enables seamless communication and integration between multiple programs.
The disclosed technology solves problems associated with bilingual or multi-lingual users receiving phone calls and voicemails in languages other than their native languages. Currently, there are no systems to deliver voicemail transcriptions in a user's preferred language. This can be an issue for busy users who can speak another language fluently but have limited capability of reading in that same language. This can lead to communication problems, which can cause the user to not understand important messages. Additionally, by transcribing and translating voicemails, users who are busy in a conference call, meeting or are in a public place where they cannot play received voicemail, can easily read the voicemail transcript received in their chosen translated language and are able to protect their privacy. For example, a user can receive a voicemail about medical test results. By receiving a transcribed message of the voicemail from a doctor's office, the user is able to tell whether they need to return the call immediately or not. This technology thus improves the user's ability to communicate with others quickly and efficiently.
The description and associated drawings are illustrative examples and are not to be construed as limiting. This disclosure provides certain details for a thorough understanding and enabling description of these examples. One skilled in the relevant technology will understand, however, that the invention can be practiced without many of these details. Likewise, one skilled in the relevant technology will understand that the invention can include well-known structures or features that are not shown or described in detail, to avoid unnecessarily obscuring the descriptions of examples.
1 FIG. 100 100 100 102 1 102 4 102 102 100 is a block diagram that illustrates a wireless telecommunication network(“network”) in which aspects of the disclosed technology are incorporated. The networkincludes base stations-through-(also referred to individually as “base station” or collectively as “base stations”). A base station is a type of network access node (NAN) that can also be referred to as a cell site, a base transceiver station, or a radio base station. The networkcan include any combination of NANs including an access point, radio transceiver, gNodeB (gNB), NodeB, eNodeB (eNB), Home NodeB or Home eNodeB, or the like. In addition to being a wireless wide area network (WWAN) base station, a NAN can be a wireless local area network (WLAN) access point, such as an Institute of Electrical and Electronics Engineers (IEEE) 802.11 access point.
100 100 104 1 104 7 104 104 106 104 100 104 102 4 The NANs of a networkformed by the networkalso include wireless devices-through-(referred to individually as “wireless device” or collectively as “wireless devices”) and a core network. The wireless devicescan correspond to or include networkentities capable of communication using various connectivity standards. For example, a 5G communication channel can use millimeter wave (mmW) access frequencies of 28 GHz or more. In some implementations, the wireless devicecan operatively couple to a base stationover a long-term evolution/long-term evolution-advanced (LTE/LTE-A) communication channel, which is referred to as aG communication channel.
106 102 106 1 104 102 106 110 1 110 3 1 The core networkprovides, manages, and controls security services, user authentication, access authorization, tracking, internet protocol (IP) connectivity, and other access, routing, or mobility functions. The base stationsinterface with the core networkthrough a first set of backhaul links (e.g., Sinterfaces) and can perform radio configuration and scheduling for communication with the wireless devicesor can operate under the control of a base station controller (not shown). In some examples, the base stationscan communicate with each other, either directly or indirectly (e.g., through the core network), over a second set of backhaul links-through-(e.g., Xinterfaces), which can be wired or wireless communication links.
102 104 112 1 112 4 112 112 112 102 100 112 The base stationscan wirelessly communicate with the wireless devicesvia one or more base station antennas. The cell sites can provide communication coverage for geographic coverage areas-through-(also referred to individually as “coverage area” or collectively as “coverage areas”). The coverage areafor a base stationcan be divided into sectors making up only a portion of the coverage area (not shown). The networkcan include base stations of different types (e.g., macro and/or small cell base stations). In some implementations, there can be overlapping coverage areasfor different service environments (e.g., Internet of Things (IoT), mobile broadband (MBB), vehicle-to-everything (V2X), machine-to-machine (M2M), machine-to-everything (M2X), ultra-reliable low-latency communication (URLLC), machine-type communication (MTC)).
100 100 102 102 100 100 102 The networkcan include a 5G networkand/or an LTE/LTE-A or other network. In an LTE/LTE-A network, the term “eNBs” is used to describe the base stations, and in 5G new radio (NR) networks, the term “gNBs” is used to describe the base stationsthat can include mmW communications. The networkcan thus form a heterogeneous networkin which different types of base stations provide coverage for various geographic regions. For example, each base stationcan provide communication coverage for a macro cell, a small cell, and/or other types of cells. As used herein, the term “cell” can relate to a base station, a carrier or component carrier associated with the base station, or a coverage area (e.g., sector) of a carrier or base station, depending on context.
100 100 100 A macro cell generally covers a relatively large geographic area (e.g., several kilometers in radius) and can allow access by wireless devices that have service subscriptions with a wireless networkservice provider. As indicated earlier, a small cell is a lower-powered base station, as compared to a macro cell, and can operate in the same or different (e.g., licensed, unlicensed) frequency bands as macro cells. Examples of small cells include pico cells, femto cells, and micro cells. In general, a pico cell can cover a relatively smaller geographic area and can allow unrestricted access by wireless devices that have service subscriptions with the networkprovider. A femto cell covers a relatively smaller geographic area (e.g., a home) and can provide restricted access by wireless devices having an association with the femto unit (e.g., wireless devices in a closed subscriber group (CSG), wireless devices for users in the home). A base station can support one or multiple (e.g., two, three, four, and the like) cells (e.g., component carriers). All fixed transceivers noted herein that can provide access to the networkare NANs, including small cells.
104 102 106 The communication networks that accommodate various disclosed examples can be packet-based networks that operate according to a layered protocol stack. In the user plane, communications at the bearer or Packet Data Convergence Protocol (PDCP) layer can be IP-based. A Radio Link Control (RLC) layer then performs packet segmentation and reassembly to communicate over logical channels. A Medium Access Control (MAC) layer can perform priority handling and multiplexing of logical channels into transport channels. The MAC layer can also use Hybrid ARQ (HARQ) to provide retransmission at the MAC layer, to improve link efficiency. In the control plane, the Radio Resource Control (RRC) protocol layer provides establishment, configuration, and maintenance of an RRC connection between a wireless deviceand the base stationsor core networksupporting radio bearers for the user plane data. At the Physical (PHY) layer, the transport channels are mapped to physical channels.
104 100 104 104 1 104 2 104 3 104 4 104 5 104 6 104 7 Wireless devices can be integrated with or embedded in other devices. As illustrated, the wireless devicesare distributed throughout the network, where each wireless devicecan be stationary or mobile. For example, wireless devices can include handheld mobile devices-and-(e.g., smartphones, portable hotspots, tablets, etc.); laptops-; wearables-; drones-; vehicles with wireless connectivity-; head-mounted displays with wireless augmented reality/virtual reality (AR/VR) connectivity-; portable gaming consoles; wireless routers, gateways, modems, and other fixed-wireless access devices; wirelessly connected sensors that provide data to a remote server over a network; IoT devices such as wirelessly connected smart home appliances; etc.
104 A wireless device (e.g., wireless devices) can be referred to as a user equipment (UE), a customer premises equipment (CPE), a mobile station, a subscriber station, a mobile unit, a subscriber unit, a wireless unit, a remote unit, a handheld mobile device, a remote device, a mobile subscriber station, a terminal equipment, an access terminal, a mobile terminal, a wireless terminal, a remote terminal, a handset, a mobile client, a client, or the like.
100 100 A wireless device can communicate with various types of base stations and networkequipment at the edge of a networkincluding macro eNBs/gNBs, small cell eNBs/gNBs, relay base stations, and the like. A wireless device can also communicate with other wireless devices either within or outside the same coverage area of a base station via device-to-device (D2D) communications.
114 1 114 9 114 114 100 104 102 102 104 114 114 114 The communication links-through-(also referred to individually as “communication link” or collectively as “communication links”) shown in networkinclude uplink (UL) transmissions from a wireless deviceto a base stationand/or downlink (DL) transmissions from a base stationto a wireless device. The downlink transmissions can also be called forward link transmissions while the uplink transmissions can also be called reverse link transmissions. Each communication linkincludes one or more carriers, where each carrier can be a signal composed of multiple sub-carriers (e.g., waveform signals of different frequencies) modulated according to the various radio technologies. Each modulated signal can be sent on a different sub-carrier and carry control information (e.g., reference signals, control channels), overhead information, user data, etc. The communication linkscan transmit bidirectional communications using frequency division duplex (FDD) (e.g., using paired spectrum resources) or time division duplex (TDD) operation (e.g., using unpaired spectrum resources). In some implementations, the communication linksinclude LTE and/or mmW communication links.
100 102 104 102 104 102 104 In some implementations of the network, the base stationsand/or the wireless devicesinclude multiple antennas for employing antenna diversity schemes to improve communication quality and reliability between base stationsand wireless devices. Additionally or alternatively, the base stationsand/or the wireless devicescan employ multiple-input, multiple-output (MIMO) techniques that can take advantage of multi-path environments to transmit multiple spatial layers carrying the same or different coded data.
100 100 116 1 116 2 100 100 100 In some examples, the networkimplements 6G technologies including increased densification or diversification of network nodes. The networkcan enable terrestrial and non-terrestrial transmissions. In this context, a Non-Terrestrial Network (NTN) is enabled by one or more satellites, such as satellites-and-, to deliver services anywhere and anytime and provide coverage in areas that are unreachable by any conventional Terrestrial Network (TN). A 6G implementation of the networkcan support terahertz (THz) communications. This can support wireless applications that demand ultrahigh quality of service (QoS) requirements and multi-terabits-per-second data transmission in the era of 6G and beyond, such as terabit-per-second backhaul systems, ultra-high-definition content streaming among mobile devices, AR/VR, and wireless high-bandwidth secure communications. In another example of 6G, the networkcan implement a converged Radio Access Network (RAN) and Core architecture to achieve Control and User Plane Separation (CUPS) and achieve extremely low user plane latency. In yet another example of 6G, the networkcan implement a converged Wi-Fi and Core architecture to increase and improve indoor coverage.
2 FIG. 200 202 202 202 204 206 208 214 216 222 226 228 230 illustrates an auto translation systemincluding a voicemail application serverconfigured to receive and manage requests to transcribe and update voicemails. The voicemail application serverincludes processing components, software, and other components configured to connect and exchange data with other devices and systems over the internet or other communications networks. As such, the voicemail application servercan connect to mobile device, voicemail profile database, voicemail message store, short message service center, message transfer agent, IP Multimedia Subsystem (IMS) network, object storage service, transcribe service, translate service, and mobile device. As such, the user can receive voicemail transcriptions in their preferred language.
202 204 232 230 234 204 230 222 232 234 222 204 230 222 202 222 202 A voicemail is a digital recorded message left by a first user (e.g., calling party) for a second user (e.g., called party) when the second user is unable to answer a phone call. It serves as a way to leave a verbal message when the recipient is unavailable or chooses not to answer the call. Voicemail messages are typically stored on the recipient's voicemailbox on a voicemail system, which is hosted by a telecommunications service provider. The voicemail application servercan receive, via the telecommunications network, a voicemail message communicated from a first user device such as mobile deviceassociated with the first user (e.g., user) to a second user device such as mobile deviceassociated with the second user (e.g., user). For example, mobile devicecan initiate a phone call to mobile device. The phone call is routed through an IMS network such as IMS networkassociated with the telecommunication network. IMS network refers to the architectural framework in the telecommunications to deliver the phone call. For instance, when the userinitiates a call to user, the IMS networkcan manage the exchange of information between the mobile deviceand mobile devicebased on Session Initiation Protocol (SIP) and Real-time Transport Protocol (RTP). This allows the telecommunication network to initiate and maintain real-time multimedia communication sessions over IP networks such as Wi-Fi and cellular networks. The IMS networkenables the delivery of voicemails to the voicemail application serverover IP-based networks. The IMS networkcan also prioritize voice traffic. This allows the voicemail application serverto receive voicemails with high-quality audio.
222 232 234 222 230 230 230 222 230 222 222 222 202 202 232 202 224 234 The IMS networkcan determine when a call from useris not accepted by user. For example, during the call, the IMS networksends an alerting signal to mobile device. Mobile devicecan transmit a response in response to the alerting signal. In one embodiment, mobile devicecan transmit a rejection signal to the IMS network. In another embodiment, mobile devicemay transmit no signal. If the IMS networkreceives no response within a certain time frame, the IMS networkdetermines the call isn't being accepted. In response to determining the call isn't being accepted, the IMS networkroutes the call to the voicemail application server. The voicemail application servercan record the voicemail message from user. The voicemail application servercan store the voicemail message as an audio file by calling a storage API to the object storage serviceif useris subscribed to the voicemail transcription service.
224 202 224 226 202 232 202 224 224 226 226 226 224 202 224 226 202 226 After recording the voicemail message and uploading it to the object storage service, the voicemail application servercan call a transcribe API that would access the voicemail message from the object storage serviceand process for transcription using transcribe service. An API is a set of protocols that allow different software applications to communicate with each other. By doing so, APIs facilitate interoperability between different systems. For example, the voicemail application servercan record a voicemail message in English, “How are you” from user. The voicemail application servercan upload the audio file into the object storage serviceusing PutObject REST API. For instance, the API can authenticate the credentials needed to access the object storage serviceto upload the voicemail audio file. The transcribe serviceis a cloud-based Automatic Speech Recognition service. The transcribe servicecan process an audio file and convert it into written text. Using ASR technology, the transcribe servicecan identify spoken words, accurate punctuations, and other linguistic elements from the voicemail message. In response to the voicemail message being uploaded to the storage, the voicemail application servercan call an API that would use a MediaFileUri (Uniform Resource Identifier) to access the audio file location on the object storage serviceto initiate a transcription of the voicemail message by the transcribe service. For example, after determining there was a successful upload, the voicemail application servercan create a transcribe job by calling StartTranscriptionJob REST API. For instance, the API can authenticate the credentials needed to access the transcribe serviceto start the transcription job.
226 202 226 202 226 202 226 202 In some embodiments, transmitting the command to initiate the transcription of the voicemail message by the transcribe service, the voicemail application servercan receive a confirmation signal from the transcribe servicevia an API. The confirmation signal indicates a successful receipt and processing of the entirety of the voicemail message. For example, the voicemail application servercan receive an API response that includes details about the transcribe job such as the job name and its status and associated timestamps. For instance, the API response may include a status as in-progress and the timestamp that the transcribe serverstarted the job. By doing so, the voicemail application servercan ensure that the transcribe serviceis functioning properly. Additionally, the voicemail application serverchecks that the whole voicemail is being transcribed.
202 226 202 226 202 202 202 In some embodiments, the voicemail application servercan invoke another API to inspect the progress of the transcription of the voicemail message by the transcribe service. The voicemail application servercan receive, via the API, an indication of an incomplete status for the transcription of the voicemail message by the transcribe service. In response to receiving the indication of the incomplete status, calling the API to inspect the progress of the transcription of the voicemail message after a time interval until the transcription of the voicemail message is completed. For example, to know the status of the transcribe job, the voicemail application servercalls GetTranscriptionJob REST API. If status of transcribe job is in-progress then the voicemail application servercalls this API again after configured time interval till the time it gets the status as completed. By doing so, the voicemail application servercan continuously check on the transcription of the voicemail message using APIs.
202 226 232 226 202 224 224 The voicemail application servercan receive, from the transcribe service, a completed transcription of the voicemail message in a default language value. For example, the default language value refers to the detected language from the audio file. For instance, if the userwas speaking in English, then the transcribe servicetranscribes the voicemail in English. For example, the voicemail application servercan retrieve the voicemail transcript from the object storage serviceusing GetObject REST API. GetObject REST API can use the TranscriptFileUri (Uniform Resource Identifier) specified in the GetTranscriptionJob API “completed” status response to access the transcript file location on the object storage service. This API allows for efficient data retrieval for large files to and from various applications.
202 206 206 234 202 206 234 The voicemail application servercan retrieve from the voicemail profile database, an indication of a target language value for the completed transcription. The voicemail profile databasecan store multiple user profiles for subscribers of the voicemail translation service. The target language value corresponds to a user-chosen language value in a user profile of the user. For example, the voicemail application servercan query voicemail profile databaseto determine the recipient (e.g., user) is a user with voicemail translation service with Spanish as a chosen translation language.
202 202 234 202 202 228 228 228 The voicemail application servercan compare the default language value of the completed transcription to the target language value. For example, the voicemail application servercan compare the language value of transcript English with userchosen translation language, Spanish. In response to determining that the default language value of the completed transcription (e.g., English) does not match the target language value (e.g., Spanish), the voicemail application servercan trigger an API to generate a translation of the completed transcription in accordance with the target language value. For example, the voicemail application servercan trigger an API request to process an English transcript to receive a Spanish text in return. The API can transmit the completed transcription to a translate service. The translate serviceis a cloud-based machine translation service. The translate servicecan use machine learning to generate the translated transcription of the voicemail. A “model,” as used herein, can refer to a construct that is trained using training data to make predictions or provide probabilities for new data items, whether or not the new data items were included in the training data. For example, training data for supervised learning can include items with various parameters and an assigned classification. A new data item can have parameters that a model can use to assign a classification to the new data item. As another example, a model can be a probability distribution resulting from the analysis of training data, such as a likelihood of an n-gram occurring in a given language based on an analysis of a large corpus from that language. Examples of models include neural networks, support vector machines, decision trees, Parzen windows, Bayes, clustering, reinforcement learning, probability distributions, decision trees, decision tree forests, and others. Models can be configured for various situations, data types, sources, and output formats.
In some implementations, the machine learning model can be a neural network with multiple input nodes that receive data inputs such as high amounts of multilingual text data. The input nodes can correspond to functions that receive the input and produce results. These results can be provided to one or more levels of intermediate nodes that each produce further results based on a combination of lower-level node results. A weighting factor can be applied to the output of each node before the result is passed to the next layer node. At a final layer (“the output layer”), one or more nodes can produce a value classifying the input that, once the model is trained, can be used to generate translations. In some implementations, such neural networks, known as deep neural networks, can have multiple layers of intermediate nodes with different configurations, can be a combination of models that receive different parts of the input and/or input from other parts of the deep neural network, or are convolutions—partially using output from previous iterations of applying the model as further input to produce results for the current input.
228 228 The translate servicecan utilize a neural machine translation (NMT) model to generate the translated transcription of the voicemail. For instance, the NMT uses artificial neural networks to translate one language to another. The translate servicecan include an encoder-decoder architecture. For example, an encoder transforms the transcription into a vector. The decoder takes the vector as an input and outputs a translated transcription of the voicemail. The decoder generates one word at a time and takes into consideration the context of the sentence and the previous words that were generated.
202 208 208 234 202 228 232 202 208 The voicemail application servercan store the translation of the completed transcription in a voicemail message store. The voicemail message storeis accessible by the user. The language value of the translation of the completed transcription matches the target language value. For example, the voicemail application servercan receive from the translate servicethe translated voicemail message transcription of “How are you” from userinto Spanish as “cómo estás”. After that the voicemail application servercan store the translated transcription in voicemail message store.
202 234 202 234 202 206 After storing the translated transcription, the voicemail application servercan generate a notification indicating a new message for the user. After that, the voicemail application servercan determine how to present the voicemail message to user. For instance, the voicemail application servercan retrieve, from the voicemail profile database, an indication of a user preference for a delivery method for notifications of translated voicemail messages.
202 202 230 208 202 214 214 214 230 In one embodiment, the voicemail application servercan determine that the user preference includes a short message service (SMS). In response to that, the voicemail application servercan transmit, to the second user device (e.g., mobile device), an SMS message including a notification about the translation of the completed transcription is stored at the voicemail message store. For example, an SMS Voicemail Message Waiting Indicator (MWI) Notification is sent by the voicemail application serverto the Short Message Service Center (SMSC). The short message service centeris an important component of the SMS infrastructure in telecommunications networks. It is responsible for managing SMS messages between mobile devices. SMSCdelivers the SMS Voicemail MWI along with the SMS containing voicemail transcription to mobile device.
202 202 212 212 212 208 234 208 216 216 230 220 In another embodiment, the voicemail application servercan determine the user preference includes an electronic mailbox. In response to determining that, the voicemail application servercan transmit a push notification to a secure push proxy server. The secure push proxy servermanages push services in telecommunication networks. The secure push proxy serverdelivers the push notification to the electronic mailbox of the second user. The push notification indicates that the translation of the completed transcription is stored at the voicemail message store. For example, in response to determining that useris registered for push notifications, the voicemail message storesends a push notification to Secure Push Proxy server. Secure Push Proxydelivers the push notification to the mobile devicevia Cloud Messaging.
202 202 208 216 216 230 218 216 218 216 218 216 234 202 216 216 218 230 In some embodiments, the voicemail application servercan determine the user preference indicates electronic mail. In response to determining that, the voicemail application servercan transmit a notification about the translation of the completed transcription stored at the voicemail message storeto a message transfer agent. The message transfer agentdelivers the notification to the second user device (e.g., mobile device) via the message delivery agent. The message transfer agentand message delivery agentare two key components in email communication systems. The message transfer agentmanages routing email messages between servers. The message delivery agentplaces the email message to the mailbox from MTA. For example, usercan set his/her email address in the voicemail profile as their preferred method of receiving messages. The voicemail application servercan send an email notification to Message Transfer Agent. The MTAdelivers the email containing the translated transcription text and attached voicemail via Message Delivery agentto the email inbox on mobile device.
230 208 210 230 234 Depending on the user preference, upon receiving the notification, mobile deviceperforms authentication and sends an API request to retrieve the voicemail message with transcription from voicemail message storevia Web Services Gateway. After receiving the voicemail message, mobile devicedisplays the voicemail message for user.
202 226 202 224 In some embodiments, the voicemail application servercan remove the uploaded message and completed transcription from the transcribe service. The voicemail application serverdeletes both the uploaded audio file and voicemail transcript from the object storage serviceusing an API to protect users'privacy.
3 FIG. 300 300 300 is a flow diagram that illustrates a methodperformed by a voicemail translation service of a telecommunications network configured to translate a voicemail message. The methodcan be performed by the system including, for example, a handheld mobile device (e.g., smartphone) and/or a server coupled to the handheld mobile device over a communications network (e.g., a telecommunications network). The handheld mobile device and/or server includes at least one hardware processor and at least one non-transitory memory storing instructions that, when executed by the at least one hardware processor, cause the system to perform the method.
302 204 232 230 234 204 230 222 At, the system can receive a voicemail message communicated from a first user device to a second user device. In one example, the system can receive receiving, via the telecommunications network, a voicemail message communicated from a first user device such as mobile deviceassociated with the first user (e.g., user) to a second user device such as mobile deviceassociated with the second user (e.g., user). For example, mobile devicecan initiate a phone call to mobile device. The phone call is routed through an IMS network such IMS networkassociated with the telecommunication network.
304 226 202 232 202 224 226 224 226 226 226 At, the system can call an API to upload the voicemail message to an object storage service. In one example, the system can call an API to upload the voicemail message to the transcribe service. An API is a set of protocols that allow different software applications to communicate with each other. By doing so, APIs facilitate interoperability between different systems. For example, the voicemail application servercan record a voicemail message in English, “How are you” from user. The voicemail application servercan upload the audio file into the object storage serviceusing PutObject REST API. The API can transmit the audio file of the voicemail message to the transcribe servicefrom the object storage service. The transcribe serviceis a cloud-based Automatic Speech Recognition service. The transcribe servicecan take an audio file and convert it into written text. Using ASR technology, the transcribe servicecan identify spoken words, accurate punctuations, and other linguistic elements from the voicemail message.
306 226 202 226 At, the system can call an API to initiate transcription of the voicemail message. The system can transmit a command to initiate a transcription of a voicemail message. In one example, the system can transmit a command to initiate a transcription of the voicemail message by the transcribe service. For example, after determining there was a successful upload, the voicemail application servercan create a transcribe job by calling StartTranscriptionJob REST API. For instance, the API can authenticate the credentials needed to access the transcribe serviceto start the transcription job.
308 226 232 226 202 224 At, the system can receive, from a transcription service, a completed transcription of the voicemail message in a default language value. In one example, the system can receive, from the transcribe service, a completed transcription of the voicemail message in a default language value. For example, the default language value refers to the detected language from the audio file. For instance, if the userwas speaking in English, then the transcribe servicetranscribes the voicemail in English. For example, the voicemail application servercan retrieve the voicemail transcript from the object storage serviceusing GetObject REST API. This API allows for efficient data retrieval for large files to and from various applications.
310 206 206 234 202 206 234 At, the system can retrieve an indication of a target language value for the completed transcription. In one example, the system can retrieve from the voicemail profile database, an indication of a target language value for the completed transcription. The voicemail profile databasecan store multiple user profiles for subscribers of the voicemail translation service. The target language value corresponds to a user-chosen language value in a user profile of the user. For example, the voicemail application servercan query voicemail profile databaseto determine the recipient (e.g., user) is a user with voicemail translation service with Spanish as a chosen translation language.
312 202 234 At, the system can compare the default language value of the completed transcription to the target language value. In one example, the system can compare the default language value of the completed transcription to the target language value. For example, the voicemail application servercan compare the language value of transcript English with userchosen translation language, Spanish.
314 202 228 At, the system can trigger another API to generate a translation of the completed transcription. In one example, the system can trigger an API to generate a translation of the completed transcription in accordance with the target language value. For example, the voicemail application servercan trigger an API request to process an English transcript to receive a Spanish text in return. The API can transmit the completed transcription to a translate service.
316 208 208 234 202 228 232 202 208 At, the system can store the translation of the completed transcription in a voicemail storage system. In one example, the system can store the translation of the completed transcription in a voicemail message store. The voicemail message storeis accessible by the user. The language value of the translation of the completed transcription matches the target language value. For example, the voicemail application servercan receive from the translate servicethe translated voicemail message transcription of “How are you” from userinto Spanish as “cómo estás”. After that the voicemail application servercan store the translated transcription in voicemail message store.
4 FIG. 4 FIG. 400 400 402 406 410 412 418 420 422 424 426 430 416 416 400 is a block diagram that illustrates an example of a computer systemin which at least some operations described herein can be implemented. As shown, the computer systemcan include: one or more processors, main memory, non-volatile memory, a network interface device, a video display device, an input/output device, a control device(e.g., keyboard and pointing device), a drive unitthat includes a machine-readable (storage) medium, and a signal generation devicethat are communicatively connected to a bus. The busrepresents one or more physical buses and/or point-to-point connections that are connected by appropriate bridges, adapters, or controllers. Various common components (e.g., cache memory) are omitted fromfor brevity. Instead, the computer systemis intended to illustrate a hardware device on which components illustrated or described relative to the examples of the figures and any other components described in this specification can be implemented.
400 400 400 400 400 The computer systemcan take any suitable physical form. For example, the computing systemcan share a similar architecture as that of a server computer, personal computer (PC), tablet computer, mobile telephone, game console, music player, wearable electronic device, network-connected (“smart”) device (e.g., a television or home assistant device), AR/VR systems (e.g., head-mounted display), or any electronic device capable of executing a set of instructions that specify action(s) to be taken by the computing system. In some implementations, the computer systemcan be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC), or a distributed system such as a mesh of computer systems, or it can include one or more cloud components in one or more networks. Where appropriate, one or more computer systemscan perform operations in real time, in near real time, or in batch mode.
412 400 414 400 400 412 The network interface deviceenables the computing systemto mediate data in a networkwith an entity that is external to the computing systemthrough any communication protocol supported by the computing systemand the external entity. Examples of the network interface deviceinclude a network adapter card, a wireless network interface card, a router, an access point, a wireless router, a switch, a multilayer switch, a protocol converter, a gateway, a bridge, a bridge router, a hub, a digital media receiver, and/or a repeater, as well as all wireless elements noted herein.
406 410 426 426 428 426 400 426 The memory (e.g., main memory, non-volatile memory, machine-readable medium) can be local, remote, or distributed. Although shown as a single medium, the machine-readable mediumcan include multiple media (e.g., a centralized/distributed database and/or associated caches and servers) that store one or more sets of instructions. The machine-readable mediumcan include any medium that is capable of storing, encoding, or carrying a set of instructions for execution by the computing system. The machine-readable mediumcan be non-transitory or comprise a non-transitory device. In this context, a non-transitory storage medium can include a device that is tangible, meaning that the device has a concrete physical form, although the device can change its physical state. Thus, for example, non-transitory refers to a device remaining tangible despite this change in state.
410 Although implementations have been described in the context of fully functioning computing devices, the various examples are capable of being distributed as a program product in a variety of forms. Examples of machine-readable storage media, machine-readable media, or computer-readable media include recordable-type media such as volatile and non-volatile memory, removable flash memory, hard disk drives, optical disks, and transmission-type media such as digital and analog communication links.
404 408 428 402 400 In general, the routines executed to implement examples herein can be implemented as part of an operating system or a specific application, component, program, object, module, or sequence of instructions (collectively referred to as “computer programs”). The computer programs typically comprise one or more instructions (e.g., instructions,,) set at various times in various memory and storage devices in computing device(s). When read and executed by the processor, the instruction(s) cause the computing systemto perform operations to execute elements involving the various aspects of the disclosure.
The terms “example,” “embodiment,” and “implementation” are used interchangeably. For example, references to “one example” or “an example” in the disclosure can be, but not necessarily are, references to the same implementation; and such references mean at least one of the implementations. The appearances of the phrase “in one example” are not necessarily all referring to the same example, nor are separate or alternative examples mutually exclusive of other examples. A feature, structure, or characteristic described in connection with an example can be included in another example of the disclosure. Moreover, various features are described that can be exhibited by some examples and not by others. Similarly, various requirements are described that can be requirements for some examples but not for other examples.
The terminology used herein should be interpreted in its broadest reasonable manner, even though it is being used in conjunction with certain specific examples of the invention. The terms used in the disclosure generally have their ordinary meanings in the relevant technical art, within the context of the disclosure, and in the specific context where each term is used. A recital of alternative language or synonyms does not exclude the use of other synonyms. Special significance should not be placed upon whether or not a term is elaborated or discussed herein. The use of highlighting has no influence on the scope and meaning of a term. Further, it will be appreciated that the same thing can be said in more than one way.
Unless the context clearly requires otherwise, throughout the description and the claims, the words “comprise,” “comprising,” and the like are to be construed in an inclusive sense, as opposed to an exclusive or exhaustive sense—that is to say, in the sense of “including, but not limited to.” As used herein, the terms “connected,” “coupled,” and any variants thereof mean any connection or coupling, either direct or indirect, between two or more elements; the coupling or connection between the elements can be physical, logical, or a combination thereof. Additionally, the words “herein,” “above,” “below,” and words of similar import can refer to this application as a whole and not to any particular portions of this application. Where context permits, words in the above Detailed Description using the singular or plural number may also include the plural or singular number, respectively. The word “or” in reference to a list of two or more items covers all of the following interpretations of the word: any of the items in the list, all of the items in the list, and any combination of the items in the list. The term “module” refers broadly to software components, firmware components, and/or hardware components.
While specific examples of technology are described above for illustrative purposes, various equivalent modifications are possible within the scope of the invention, as those skilled in the relevant art will recognize. For example, while processes or blocks are presented in a given order, alternative implementations can perform routines having steps, or employ systems having blocks, in a different order, and some processes or blocks may be deleted, moved, added, subdivided, combined, and/or modified to provide alternative or sub-combinations. Each of these processes or blocks can be implemented in a variety of different ways. Also, while processes or blocks are at times shown as being performed in series, these processes or blocks can instead be performed or implemented in parallel, or can be performed at different times. Further, any specific numbers noted herein are only examples such that alternative implementations can employ differing values or ranges.
Details of the disclosed implementations can vary considerably in specific implementations while still being encompassed by the disclosed teachings. As noted above, particular terminology used when describing features or aspects of the invention should not be taken to imply that the terminology is being redefined herein to be restricted to any specific characteristics, features, or aspects of the invention with which that terminology is associated. In general, the terms used in the following claims should not be construed to limit the invention to the specific examples disclosed herein, unless the above Detailed Description explicitly defines such terms. Accordingly, the actual scope of the invention encompasses not only the disclosed examples but also all equivalent ways of practicing or implementing the invention under the claims. Some alternative implementations can include additional elements to those implementations described above or include fewer elements.
Any patents and applications and other references noted above, and any that may be listed in accompanying filing papers, are incorporated herein by reference in their entireties, except for any subject matter disclaimers or disavowals, and except to the extent that the incorporated material is inconsistent with the express disclosure herein, in which case the language in this disclosure controls. Aspects of the invention can be modified to employ the systems, functions, and concepts of the various references described above to provide yet further implementations of the invention.
To reduce the number of claims, certain implementations are presented below in certain claim forms, but the applicant contemplates various aspects of an invention in other forms. For example, aspects of a claim can be recited in a means-plus-function form or in other forms, such as being embodied in a computer-readable medium. A claim intended to be interpreted as a means-plus-function claim will use the words “means for.” However, the use of the term “for” in any other context is not intended to invoke a similar interpretation. The applicant reserves the right to pursue such additional claim forms either in this application or in a continuing application.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
February 23, 2026
July 2, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.