Patentable/Patents/US-12694213-B2
US-12694213-B2

Entropy-based set block decoding utilizing an adaptive entropy threshold

PublishedJuly 28, 2026
Assigneenot available in USPTO data we have
Technical Abstract

Certain aspects of the disclosure provide techniques for language model token prediction. A method generally includes processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position; and an entropy score based on the probability distribution; for each respective token position, determining whether the entropy score satisfies a current adaptive entropy threshold, which is based on the entropy score associated with each respective token position; based on the determination for each respective token position, performing a decoding operation comprising: a set block decoding operation or a single-token decoding operation to generate one or more output tokens that correspond to one or more token positions; and generating a first version of the output token sequence including the one or more output tokens.

Patent Claims

Legal claims defining the scope of protection, as filed with the USPTO.

1

a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens; processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, determining whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions; a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence; or a single-token decoding operation to generate a single output token that corresponds to the first token position; and based on the determination for each respective token position of the set of token positions, performing a decoding operation comprising: generating a first version of the output token sequence including the at least two output tokens or the single output token. . A method of token prediction, comprising:

2

claim 1 the set of token positions comprises at least two token positions, of the plurality of token positions, that correspond to at least the first token position and a second token position in the output token sequence, the decoding operation comprises the set block decoding operation based on the entropy score associated with each respective token position of at least two token positions satisfying the current adaptive entropy threshold, and generating the first version of the output token sequence comprises generating the first version of the output token sequence with the at least two output tokens. . The method of, wherein:

3

claim 2 . The method of, wherein performing the set block decoding operation comprises decoding, in parallel, the sets of candidate output tokens corresponding to the at least two token positions to generate the at least two output tokens.

4

claim 1 the set of token positions comprises the first token position and the second token position in the output token sequence, the entropy score associated with the first token position satisfying the current adaptive entropy threshold; and the entropy score associated with the second token position not satisfying the current adaptive entropy threshold, and the decoding operation comprises the single-token decoding operation based on: generating the first version of the output token sequence comprises generating the first version of the output token sequence with the single output token. . The method of, wherein:

5

claim 1 the set of token positions comprises the first token position in the output token sequence, the decoding operation comprises the single-token decoding operation based on the entropy score associated with the first token position not satisfying the current adaptive entropy threshold, and generating the first version of the output token sequence comprises generating the first version of the output token sequence with the single output token. . The method of, wherein:

6

claim 1 determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions; updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores; determining one or more trend metrics based on the sliding window of entropy scores; and determining the current adaptive entropy threshold based on the trend metrics. . The method of, further comprising:

7

claim 6 a mean entropy score; a median entropy score; or a weighted aggregation entropy score. . The method of, wherein the current decoding cycle entropy score comprises:

8

claim 6 a slope; or a variance. . The method of, wherein the one or more trend metrics comprise at least one of:

9

claim 6 . The method of, wherein determining the current adaptive entropy threshold comprises determining a first mapping between the one or more trend metrics and the current adaptive entropy threshold.

10

claim 6 . The method of, wherein determining the current adaptive entropy threshold comprises determining one or more rules.

11

claim 6 determining at least one of the one or more trend metrics satisfies at least one threshold; and adjusting, based on the determination, a previous adaptive entropy threshold by an increment. . The method of, wherein determining the current adaptive entropy threshold comprises:

12

claim 11 the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is greater than a slope threshold, and adjusting the previous adaptive entropy threshold comprises decreasing the previous adaptive entropy threshold by the increment. . The method of, wherein:

13

claim 11 the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is less than a slope threshold, and adjusting the previous adaptive entropy threshold comprises increasing the previous adaptive entropy threshold by the increment. . The method of, wherein:

14

claim 6 the method further comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and determining the current adaptive entropy threshold comprises adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback. . The method of, wherein:

15

claim 14 the plurality of previous decoding cycle entropy scores; a latency metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles; or an accuracy metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles. . The method of, wherein the decoding feedback comprises information about at least one of:

16

a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens; processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions; updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores; determining one or more trend metrics based on the sliding window of entropy scores; determining a current adaptive entropy threshold based on the trend metrics; and predicting one or more tokens, associated with one or more token positions in the plurality of token positions, in the output token sequence based on the current adaptive entropy threshold. . A method of adaptive entropy threshold adjustment for token prediction, comprising:

17

claim 16 a mean entropy score; a median entropy score; or a weighted aggregation entropy score. . The method of, wherein the current decoding cycle entropy score comprises:

18

claim 16 a slope; or a variance. . The method of, wherein the one or more trend metrics comprise at least one of:

19

claim 16 the method further comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and determining the current adaptive entropy threshold comprises adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback. . The method of, wherein:

20

a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens; process an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, determine whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions; a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence; or a single-token decoding operation to generate a single output token that corresponds to the first token position; and based on the determination for each respective token position of the set of token positions, perform a decoding operation comprising: generate a first version of the output token sequence including the at least two output tokens or the single output token. . A processing system, comprising: memory comprising computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to:

Detailed Description

Complete technical specification and implementation details from the patent document.

Aspects of the present disclosure relate to language model-based token prediction.

Recent advances in artificial intelligence (AI) have enabled the widespread adoption of language models for the performance of natural language processing (NLP) tasks, such as generating text (e.g., synthetic text), answering user queries in a conversational manner, translating text from one language to another, and/or the like. A “language model” is a type of machine learning (ML) model trained on large volumes of text to learn the structure, meaning, and usage patterns of language. Language models make it possible for software to “understand” typical human speech or written content and respond to it by, in some cases, generating human-understandable responses through natural language generation (NLG).

Generating text with a language model is an autoregressive process that involves the language model predicting a next token given an input sequence of tokens (e.g., in some cases, including token(s) that were previously predicted by the language model). More specifically, given an input sequence of tokens, a language model may assign probabilities to each candidate output token in its vocabulary, where a probability assigned to a candidate output token represents a likelihood that the candidate output token is most likely to logically follow next in the sequence. The language model may select the next token by sampling according to these probabilities. For example, the language model may select, as the next token in the sequence, a candidate output token associated with a greatest probability (e.g., indicating that the token is the most likely and appropriate next token in the sequence).

In the context of language models, “tokens” may refer to units of text that the models process and generate. Tokens can represent individual characters, words, subwords, or even larger linguistic units, depending on the specific tokenization (e.g., segmentation of text into meaningful units to capture its semantic and syntactic structure) approach used. Tokens act as a bridge between text data and the numerical representations with which language models are able to use. A “candidate output token” refers to a token that may be generated by a language model as a potential next token in a text sequence.

Certain aspects provide a method of token prediction. The method includes processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens an entropy score based on the probability distribution for the set of candidate output tokens; determining, for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions; performing, based on the determination for each respective token position of the set of token positions, a decoding operation comprising: a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence a single-token decoding operation to generate a single output token that corresponds to the first token position; and generating a first version of the output token sequence including the at least two output tokens or the single output token.

Certain aspects provide a method of adaptive entropy threshold adjustment for token prediction. The method includes processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each candidate output token in the set of candidate output tokens an entropy score based on the probability distribution for the set of candidate output tokens; determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions; updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores; determining one or more trend metrics based on the sliding window of entropy scores; determining a current adaptive entropy threshold based on the trend metrics; and predicting one or more tokens, associated with one or more token positions in the plurality of token positions, in the output token sequence based on the current adaptive entropy threshold.

Other aspects provide processing systems configured to perform the aforementioned methods as well as those described herein; non-transitory, computer-readable media comprising instructions that, when executed by a processors of a processing system, cause the processing system to perform the aforementioned methods as well as those described herein; a computer program product embodied on a computer readable storage medium comprising code for performing the aforementioned methods as well as those further described herein; and a processing system comprising means for performing the aforementioned methods as well as those further described herein.

The following description and the related drawings set forth in detail certain illustrative features of one or more aspects.

To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the drawings. It is contemplated that elements and features of one embodiment may be beneficially incorporated in other embodiments without further recitation.

Language models conventionally generate text one token at a time in an autoregressive manner, where each newly generated token is conditioned on one or more previously generated tokens. This next-token prediction paradigm may deliver high-quality results by maintaining the causal relationship between tokens. However, such prediction may be computationally expensive as each token generation step may involve performing a forward pass through a language model (e.g., resulting in repeated memory accesses to obtain cached values), making the process slow and resource intensive. Set block decoding (SBD) is one approach that aims to address these inefficiencies, while also preserving accuracy.

Set block decoding refers to a method that may be used to accelerate inference in language models by enabling parallel generation of multiple future tokens. For example, token generation may occur in blocks rather than strictly token-by-token. As used herein, a “block” may refer to a contiguous sequence of tokens processed together during inferencing (e.g., output generation) by a language model. In certain aspects, set block decoding introduces bidirectional masked attention within blocks of future tokens while preserving autoregressive attention for past tokens. Autoregressive attention helps to ensure that each token is influenced by prior token(s), maintaining the causal relationship necessary for sequential prediction. Bidirectional masked attention allows attention to be applied in both forward and backward directions within a sequence, while masking certain tokens to prevent them from influencing predictions. Bidirectional masked attention may enable simultaneous processing of visible tokens and help to ensure contextual relevance during inference. Set block decoding may significantly reduce computational overhead, in some cases achieving up to three to five times fewer forward passes during inferencing, without requiring architectural changes and/or additional parameters.

A key technical challenge of set block decoding lies in identifying which tokens can be decoded in parallel without compromising the accuracy and coherence of the generated tokens. As such, a language model may need to evaluate the dependencies between tokens within a block to ensure that parallel decoding does not violate syntactic or semantic relationships. For example, certain tokens may rely heavily on the context provided by preceding tokens, making them unsuitable for parallel decoding. Addressing this challenge may involve implementing mechanism(s) to dynamically assess token uncertainty and interdependencies, such as to help ensure that parallel decoding maintains the integrity of the output.

One example mechanism that quantifies uncertainty at each future token position (simply referred to herein as “token position”) uses an entropy measure and an entropy threshold to regulate parallel decoding. Specifically, for each respective future token position, a language model may generate a probability distribution over candidate output tokens associated with the respective future token position and compute an entropy score from this probability distribution. Future token positions associated with an entropy score that satisfy the entropy threshold (e.g., entropy score < entropy threshold), indicating low entropy, may be prioritized for parallel decoding. Conversely, future token positions associated with an entropy score that do not satisfy the entropy threshold (e.g., entropy score ≥ entropy threshold), indicating high entropy, may be better suited for conservative single-token decoding.

In this context, an “entropy score” may serve as a measure of the language model's uncertainty (e.g., entropy) about which candidate output token (e.g., in its vocabulary) it may generate for a future token position. A “high entropy score” may indicate low language model confidence (e.g., high entropy) with respect to its output and represent a measure for a more uniform probability distribution over candidate output tokens. For example, high entropy may occur when a language model is prompted with an open-ended question such as, “my favorite fruit is,” where many candidate output tokens are plausible and the language model's uncertainty is high. A “low entropy score,” on the other hand, may reflect high language model confidence (e.g., low entropy) and reflect a measure for a sharply peaked probability distribution over candidate output tokens. Low entropy is typical in deterministic token generation scenarios (e.g., including factual statements, common phrases, etc.), such as, “the capital of France is,” where the language model is highly confident that “Paris” is the correct next token in the sequence.

For conventional set block decoding, the entropy threshold may be static (e.g., fixed prior to inference) and uniformly applied to determine which token positions qualify for parallel decoding. While straightforward to implement, the use of a static entropy threshold imposes a rigid speed versus accuracy trade-off that often fails to adapt to the dynamic uncertainty profiles encountered during inferencing. For example, a static entropy threshold that is set too low may restrict parallel decoding to only the most confident token positions, preserving accuracy but significantly reducing speed benefits, often approaching the behavior of next-token decoding. Conversely, a static entropy threshold set too high may permit parallel decoding for many token positions even under elevated uncertainty, increasing the risk of inaccurate token predictions that propagate errors, degrade coherence, and/or necessitate costly corrections.

Moreover, an optimal entropy threshold may be sensitive to the nature of the task and/or context. For example, fluent narrative text may tolerate aggressive parallelization, whereas mathematical reasoning and/or syntax-constrained code generation may demand stricter gating (e.g., using a lower entropy threshold to control which future token positions are allowed to be decoded in parallel). Further, rapid context shifts, such as transitioning from descriptive text to numeric calculations, may render a static entropy threshold ineffective. Attempting to manually tune the entropy threshold for specific tasks, language models, and/or deployments may introduce operational complexity, inconsistent performance, and/or poor generalization across domains, undermining the potential of set block decoding for fast and reliable inference.

Aspects described herein overcome the aforementioned technical problems and improve upon the state of the art by introducing techniques for entropy-based set block decoding utilizing an adaptive entropy threshold. Instead of relying on a static entropy threshold, the techniques described herein may be used to dynamically adjust an adaptive entropy threshold, used for set block decoding at each decoding step, based on the evolving uncertainty profile of a language model during inferencing. As used herein, a “decoding step,” or “decoding cycle,” may refer to a single iteration or operation during the inferencing process of a language model, where the language model generates or predicts one or more next tokens in a sequence. For example, in certain aspects, entropy trends of a language model may be analyzed by evaluating changes in language model uncertainty over successive decoding steps and deriving trend metrics, such as slope and/or variance, to inform threshold adjustments. When the trend metrics indicate stability and low language model uncertainty, the adaptive entropy threshold may be increased to allow more token positions to be decoded in parallel, improving efficiency. Conversely, when the trend metrics signal rising language model uncertainty, the adaptive threshold may be decreased to restrict parallel decoding, ensuring accuracy and coherence in the generated outputs. For example, in fluent text generation, the adaptive entropy threshold may be elevated for common words and punctuation, while in syntax-sensitive code generation, the adaptive entropy threshold may be lowered to avoid errors. This adaptive entropy threshold adjustment may occur in real-time, ensuring consistent decision-making across decoding steps. As used herein, “real-time” may refer to the immediate or near-instantaneous adjustment of an entropy threshold during each decoding cycle, prior to performing set block decoding or single-token decoding for the specific decoding cycle, such as to ensure consistent decision-making and token prediction without delays.

3 FIG. A system capable of performing entropy-based set block decoding, utilizing an adaptive entropy threshold, may include an entropy monitor to track language model uncertainty (e.g., token-level uncertainty), an adaptive scheduler to modulate the adaptive entropy threshold, and an entropy-bounded sampler to select tokens for decoding, such as a block of tokens in multiple token positions for parallel decoding. An example system including an entropy monitor, an adaptive schedule, and an entropy-bounded sampler is depicted and described with respect to. This system eliminates the need for manual threshold tuning and balances decoding speed with reliability across diverse applications, such as conversational AI, precise mathematical computations, and/or code generation.

The techniques described herein provide notable technical advantages over conventional solutions, such as improved efficiency in language model inferencing and the ability to dynamically adjust decoding thresholds based on entropy trends. By leveraging adaptive entropy thresholds, the techniques presented herein address technical challenges associated with static threshold approaches, such as suboptimal decoding speed and reduced accuracy in uncertain contexts. For example, the techniques described herein enable real-time analysis of entropy trends during inference, allowing for the dynamic adjustment of entropy thresholds to balance efficiency and accuracy. This functionality facilitates faster decoding in stable contexts while preserving coherence and fidelity in unstable contexts. The technical effects of these techniques enable language models to operate more efficiently, reducing computational overhead and improving scalability in high-demand environments.

Notably, the adaptive entropy threshold techniques described herein can further improve the function of any existing application that relies on language model inference for generating outputs. For example, any application that processes natural language inputs, generates predictive text, or performs translation tasks can benefit from the dynamic adjustment capabilities provided by these techniques. In some cases, the entropy threshold can be adjusted dynamically during inferencing, allowing the application to optimize decoding efficiency while maintaining accuracy. For instance, when the system detects low uncertainty in token predictions, the threshold is increased to enable faster parallel decoding. Conversely, when uncertainty rises, the threshold is decreased to ensure coherence and fidelity in the generated outputs. These technical effects allow applications to deliver faster and more reliable results, enhancing user experience and reducing computational demands.

A language model implementing these techniques, described herein, integrates adaptive entropy threshold adjustments directly into its inferencing process. This dynamic capability allows the language model to analyze entropy trends in real-time and optimize decoding efficiency and accuracy based on evolving uncertainty profiles. By departing from static entropy threshold approaches, the language model exhibits unique functionality that addresses technical challenges inherent in conventional systems, ensuring improved performance and reliability across diverse applications.

1 FIG. 1 FIG. 100 104 100 150 150 102 102 120 150 102 120 depicts an example systemsupporting a plurality of microservices(e.g., software-defined services, which in some cases, may be cloud-native). As shown in, systemincludes one or more client devices(collectively referred to herein as “client devices”) and one or more hosts(collectively referred to herein as “hosts”). A networkmay provide connectivity between client deviceand host. Networkmay include, for example, a direct link, a local area network (LAN), a wide area network (WAN) (such as the Internet), another type of network, or a combination of one or more of these networks.

102 102 102 106 106 102 Hostmay be geographically co-located servers on the same rack or on different racks in any arbitrary location in a data center. Hostmay be implemented on a server-grade hardware platform. Hostor the hardware platform may include components of a computing device, such as one or more processors (e.g., central processing units (CPUs)), one or more memories (e.g., random access memory (RAM)), one or more network interfaces (e.g., physical network interfaces (PNICs)), storage, and/or other components, as described elsewhere herein. Storageand other example components of an apparatus that may implement hostare described elsewhere herein.

102 100 104 104 104 102 102 102 104 104 104 104 104 Hostin systemmay host a set of one or more microservices(collectively referred to herein as “microservice(s)”). The microservice(s)may be deployed using virtual machines (VMs) and/or container(s) implemented on host). For example, hostmay implement a hypervisor (not shown) that abstracts processor, memory, storage, and networking resources of host's hardware platform). Generally, a microserviceis a loosely coupled and independently deployable service or software that, alone or in combination with one or more other microservices, may make up an application. Microservice(s)may enable segmented, granular level functionalities within a larger system infrastructure. A reference to a single microservicecan encompass multiple microservices, unless context indicates otherwise.

150 152 152 104 120 150 104 150 104 120 Client devicemay include a user interface (UI). UImay be usable to communicate with microservicevia network. For example, communication between client devicesand a microservicemay be facilitated by one or more application programming interfaces (APIs). An API is a set of rules and protocols that allows different software applications to communicate and share data with each other. Non-exhaustive examples of client devicesmay include a smartphone, a personal computer, a tablet, or a laptop computer. In some examples, microservicemay interact with another microservice, an application, a host, or the like, via network.

1 FIG. 104 120 108 104 102 104 104 102 104 As shown in, in certain aspects, microserviceimplements an entropy-based set block decoding service, which is any networkaccessible service that dynamically adjusts entropy thresholds during decoding to optimize efficiency and accuracy. The entropy-based set block decoding service may include a language model. The entropy-based set block decoding service may analyze uncertainty profiles and entropy trends of a language model to determine optimal adaptive entropy thresholds for decoding, such as to allow for, in some cases, the parallel decoding of multiple tokens (e.g., a block of tokens corresponding to two more future token positions), enabling faster and more reliable language model outputs. A microservice, or a hostthat implements a microservice, may be referred to as an apparatus.” A microservice, or a hostthat implements a microservice, may be referred to as an apparatus.

1 FIG. 1 FIG. 102 106 150 102 106 150 102 150 102 150 150 104 102 104 Thoughdepicts host, storage, and client deviceas single devices for ease of illustration, host, storage, and/or client devicemay be embodied in a variety of forms. Further, thoughdepicts only one hostand one client device, other examples may include a different number of hostsand/or client devices. Client devicesmay use any combination of microserviceson any hostwhere microservicesare deployed.

2 FIG. 200 200 202 depicts an example systemfor token prediction using set block decoding with an adaptive entropy threshold. Systemis configured to dynamically adjust an adaptive entropy threshold, such as to optimize token generation accuracy and speed during token prediction for an input token sequence.

2 FIG. 200 206 210 214 218 222 226 206 226 204 204 As shown in, systemincludes a token generator, an entropy scores generator, an entropy monitor, an adaptive entropy scheduler, an entropy-based sampler, and a token predictorto perform token prediction. Further, as shown, token generatorand token predictormay implement a language model. In certain aspects, the language modelmay comprise a so-called large language model (LLM). An LLM is a type of language model that has a large number of parameters, such a language model with greater than 100 billion parameters (although, it is noted, that the number of parameters generally associated with an LLM may change over time).

202 200 204 202 202 200 204 202 202 202 200 204 204 152 1 152 2 202 1 FIG. Input token sequencemay represent a user query, a prompt, and/or any sequence of token(s) provided to system(and more specifically, language model) for processing. For example, input token sequencemay comprise a question, such as “what is the capital of Texas?,” a command or instructions, such as “tell me my favorite tropical fruit,” or a partially completed sentence, such as “the tallest mountain above sea level is.” The input token sequencemay be provided to systemto generate token(s) (e.g., output token(s)) (e.g., via language model) in response to the input token sequence. In certain aspects, the input token sequenceincludes only tokens provided by a user. In certain aspects, the input token sequenceincludes tokens provided by a user and one or more tokens previously generated by system, via language model, during a previous decoding cycle. For example, the language modelmay receive, as input (e.g., such as from a user via a user interface, such as UI() and/or UI() in), tokens “the tallest mountain” and generate tokens “above,” “sea,” “level,” and “is” during previous decoding cycle(s), such that for a subsequent decoding cycle the input token sequenceincludes tokens “the tallest mountain above sea level is.”

206 204 202 208 204 202 202 204 208 202 208 202 208 208 Token generator, utilizing language model, processes the input token sequenceto generate probability distributionsfor candidate output tokens at multiple future token positions. For example, language modelmay process the input token sequence, during token generation, by partitioning the input token sequenceinto individual tokens and analyzing their context to predict the most appropriate next tokens (e.g., corresponding to future token positions) in the token sequence. For each token position, language modelmay generate a probability distributionfor a vocabulary of candidate output tokens. For example, when processing an input token sequence, language model may generate (1) a first probability distributionfor the vocabulary of candidate output tokens in a first token position after the input token sequence, (2) a second probability distributionfor the vocabulary of candidate output tokens in a second token position after the input token sequence and after the first token position, (3) a third probability distributionfor the vocabulary of candidate output tokens in a third token position after the input token sequence, the first token position, and the second token position, and so on (e.g., [input token sequence][first token position][second token position][third token position] etc.).

204 204 202 208 204 202 202 The vocabulary of candidate output tokens may include a finite, defined set of tokens learned by the language modelduring training, and which may be generated by language modelbased on input token sequence. A probability distribution, generated for a future token position, may be based on a probability score assigned, by language model, to each candidate output token in the vocabulary for that particular future token position. A probability score assigned to a specific candidate output token may represent a likelihood or confidence that the particular candidate output token is a most appropriate token at the particular future token position, given the input token sequence(and/or token(s) of one or more prior token positions). Thus, the probability score assigned to the specific candidate output token, for the particular future token position, may reflect both the syntactic and semantic relevance of the particular candidate output token to the input token sequence. A candidate output token with a probability score higher than all other candidate tokens for a specific future token position may be considered the token that the language model predicts as the most appropriate or accurate representation for that position based on the input token sequence.

LM <t <t <t 1 t-1 202 202 202 204 In certain aspects, the probability score assigned to a particular candidate output token x (e.g., for a particular future token position) may be represented as P(x|x), where xrepresents the input token sequence. For example, in certain aspects, input token sequencemay be represented as x=x, . . . , x. The input token sequencemay include t−1 tokens from a vocabulary, V, of language model.

210 208 212 206 208 210 212 208 210 204 202 204 204 204 204 Entropy scores generatorprocesses the probability distributionsgenerated for each future token position to generate entropy scores. For instance, in an example case where token generatorgenerates three probability distributionsfor three future token positions, entropy scores generatormay generate three entropy scoresbased on the three probability distributions, one for each future token position. An “entropy score” generated by entropy scores generatorfor a future token position may serve as a measure of language model's uncertainty about which candidate output token (e.g., in its vocabulary) it may generate for the future token position, such as based on the input token sequence. A high entropy score, determined for a future token position, may indicate low language modelconfidence with respect to its output for that future token position (e.g., language modelhas assigned similar probability scores to many of its candidate output tokens for that future token position). On the other hand, a low entropy score, determined for a future token position, may indicate high language modelconfidence with respect to its output for that future token position (e.g., language modelhas assigned a probability score to one of its candidate output tokens that is significantly greater than the probability scores assigned to its other candidate output tokens for that future token position).

212 210 208 212 212 In certain aspects, entropy scoresmay be generated by entropy scores generatorby applying Shannon's entropy formula to probability distributionsto generate the entropy scores(He). For example, entropy scores generator may apply Shannon's entropy formula to a probability distribution for each future token position to generate an entropy scorefor each future token position. Shannon's entropy formula may be given as:

LM <t where P(x|x) represents the probability score assigned to a particular candidate output token x, at a particular future token position, as described above.

214 212 214 212 214 212 214 212 Entropy monitormay utilize entropy scores, generated for the future token positions, to calculate a current decoding cycle entropy score. The “current decoding cycle entropy score” may refer to a single aggregated entropy score calculated for the specific decoding cycle. In certain aspects, entropy monitormay compute the current decoding cycle entropy score as a mean of entropy scores. In certain aspects, entropy monitormay compute the current decoding cycle entropy score as a median of entropy scores. In certain aspects, entropy monitormay compute the current decoding cycle entropy score by computing a weighted aggregation of the entropy scores.

214 214 216 204 216 Entropy monitormay add the current decoding cycle entropy score to a sliding window, which tracks previous decoding cycle entropy scores. As used herein, a “sliding window” may refer to a fixed-length sliding window of size W (e.g., W=4), updated each decoding cycle to include the current decoding cycle entropy score and a plurality of immediately preceding decoding cycle entropy scores, from which trend metrics are computed. For example, following the addition, an example sliding window may include [Previous decoding cycle (x−3) entropy score, previous decoding cycle (x−2) entropy score, previous decoding cycle (x−1) entropy score, current decoding cycle (x) entropy score]. Using the values in the sliding window, the entropy monitormay determine trend metrics, such as slope and/or variance, which reflect changes in language modeluncertainty over time. As described below, these trend metricsmay be used to inform (1) selection of an adaptive entropy threshold or (2) adjustment(s) to a previous adaptive entropy threshold (e.g., used for a previous decoding cycle), such as to help achieve efficient and accurate token prediction.

218 218 220 216 204 218 220 216 220 Adaptive entropy scheduler(simply referred to herein as “scheduler”) may dynamically select and/or adjust an adaptive entropy threshold for the current decoding cycle (referred to herein as “current adaptive entropy threshold”) based on trend metrics, such as to optimize performance and efficiency in decoding by language model. For example, in certain aspects, schedulerdetermines the current adaptive entropy thresholdbased on a mapping (e.g., an existing mapping) between trend metricsand the current adaptive entropy threshold, enabling precise threshold selection in response to observed trends.

218 220 218 216 216 204 204 220 220 216 220 220 In certain aspects, schedulerdetermines the current adaptive entropy thresholdbased on one or more rules. In certain aspects, schedulerevaluates whether one or more of the trend metrics(e.g., the slope and/or variance) satisfy one or more thresholds (e.g., a slope threshold and/or a variance threshold, respectively) and adjusts a previous entropy threshold (e.g., for a previous decoding cycle) by an increment, when the trend metric(s) satisfy the threshold(s). For example, in some cases where trend metricsincludes a slope that exceeds a slope threshold (e.g., represents stable entropy for the language model, indicating that language modelconfidence is stable), the adjustment may include decreasing the previous adaptive entropy threshold. Thus, the current adaptive entropy thresholdmay be a value that is less than a value of the previous adaptive entropy threshold. This current adaptive entropy thresholdmay allow more token positions to qualify for set block decoding and thus be decoded in parallel (e.g., allow for more parallelization). Alternatively, in some cases where the trend metricsincludes a slope that exceeds the slope threshold (e.g., represents a spike in entropy, indicating that model confidence is decreasing or model uncertainty is increasing), the adjustment may include increasing the previous adaptive entropy threshold. Thus, the current adaptive entropy thresholdmay be a value that is less than a value of the previous adaptive entropy threshold. This current adaptive entropy thresholdmay allow fewer token positions to qualify for set block decoding and thus be decoded in parallel.

218 220 220 In certain aspects, schedulerincorporates decoding feedback from previous cycles, including entropy scores, latency metrics, and/or accuracy metrics, to select the current adaptive entropy thresholdand/or adjust a previous adaptive entropy threshold to determine the current adaptive entropy threshold. Each metric may influence the adaptive entropy threshold in a different way. For example, where the metric is entropy trend (e.g., mean, median, and/or slope), and entropy is rising (e.g., model is becoming more uncertain), the adaptive entropy threshold may be reduced, making decoding more conservative. As another example, where the metric is latency, and the latency indicates that recent decoding cycles are too slow, the adaptive entropy threshold may be increased, such as to expand parallel decoding. Alternatively, if the latency indicates that decoding cycles are efficient/quick enough, then no upward adjustment of the adaptive entropy threshold may be needed. As another example, where the metric is accuracy, and accuracy is decreasing (e.g., errors are increasing), the adaptive entropy threshold may be decreased to reduce aggressive parallel sampling and decoding. Alternatively, if the accuracy remains high, then the adaptive entropy threshold may be allowed to drift upward (e.g., gradually increased). These metrics may bias the adaptive entropy threshold towards speed (e.g., a higher threshold) or accuracy (e.g., a lower threshold), depending on recent decoding behavior.

218 This feedback-driven approach may utilize reinforcement learning techniques, a method in which an agent learns to make decisions by receiving rewards and/or penalties for actions taken in an environment. By leveraging reinforcement learning, the schedulermay continuously refine its adjustment(s) and/or selection of adaptive entropy thresholds for decoding cycles, helping to achieve adaptability and/or improvement in decoding operations.

Example rewards may be received (1) when an increased adaptive entropy threshold leads to correct multi-token predictions (e.g., fast and accurate), (2) when latency decreases without negatively affecting output quality, and/or the like. Example penalties may be received (1) when an aggressive adaptive entropy threshold causes inaccurate token prediction and/or causes the need for fallback corrections, (2) when latency increases because the adaptive entropy threshold was set to be too low (e.g., overly conservative), and/or the like. These signals may enable a reinforcement learning controller to learn whether increasing and/or decreasing the adaptive entropy threshold improves overall speed-accuracy performance.

222 222 212 220 222 212 220 212 220 212 220 Entropy-based sampler(simply referred to herein as “sampler”) may evaluate the entropy score, determined for each future token position, individually against the current adaptive entropy thresholdto identify future token position(s) that are eligible for parallel decoding. For example, samplermay compare the entropy scoredetermined for a first future token position to the current adaptive entropy threshold, then move to the next future token position and perform the same comparison, continuing this process for a set (e.g. one or more) of future token positions. Token positions having entropy scoresthat satisfy the current adaptive entropy thresholdmay be decoded in parallel, while those that do not, may be decoded individually and sequentially (e.g., via multiple decoding cycles). In certain aspects, contiguous token positions, having entropy scoresthat satisfy the current adaptive entropy threshold, be decoded in parallel.

222 212 202 220 212 220 212 220 222 For example, in certain aspects, samplerdetermines whether the entropy scorefor the first token position (e.g., next token position after the input token sequence) satisfies the current adaptive entropy threshold. In a first case, the entropy scorefor the first token position may not satisfy the current adaptive entropy threshold(e.g., entropy score> current adaptive entropy threshold); thus, samplermay determine that a single-token decoding operation may be performed to decode a token for the first token position only (e.g., during this decoding cycle).

212 220 212 220 222 212 220 212 220 212 220 222 In a second case, the entropy scorefor the first token position may satisfy the current adaptive entropy threshold(e.g., entropy score< current adaptive entropy threshold); thus, samplermay determine whether the entropy scorefor the second token position (e.g., next token position after the first token position) satisfies the current adaptive entropy threshold. The entropy scorefor the second token position may not satisfy the current adaptive entropy threshold(e.g., entropy score> current adaptive entropy threshold); thus, samplermay determine that a single-token decoding operation may be performed to decode a token for the first token position only (e.g., during this decoding cycle).

212 220 212 220 222 212 220 212 220 212 220 222 212 220 In a third case, the entropy scorefor the first token position may satisfy the current adaptive entropy threshold(e.g., entropy score< current adaptive entropy threshold); thus, samplermay determine whether the entropy scorefor the second token position (e.g., next token position after the first token position) satisfies the current adaptive entropy threshold. The entropy scorefor the second token position may also satisfy the current adaptive entropy threshold(e.g., entropy score< current adaptive entropy threshold); thus, samplermay determine that a set block decoding operation may be performed to decode tokens for at least the first and second token positions. Other token positions with entropy scoresthat satisfy the current adaptive entropy thresholdmay also be decoded in parallel during the set block decoding operation.

222 224 226 226 224 224 226 224 226 226 204 The samplerprovides an indication of token position(s)that can be decoded together to token predictor. Token predictormay use the indication of token position(s)to generate token(s) for the identified position(s). For example, in certain aspects where the indication of token position(s)indicates only the first future token position, token predictormay perform a single-token decoding operation to generate a single output token that corresponds to the first future token position. In certain aspects where the indication of token position(s)indicates at least the first and second future token positions, token predictormay perform a set block decoding operation to generate at least two output tokens that correspond to at least the first future token position and the second future token position (and output token(s) for one or more other indicated token position(s)). Token predictormay utilize language modelto generate the output token(s) for the indicated future token position(s).

226 204 202 230 226 204 230 In certain aspects, the output token(s), generated by token predictor(e.g., via language model), may be combined and appended with the input token sequenceto generate an output token sequence. This may occur in cases where the output token(s) generated by token predictor(e.g., via language model) cause the entire sequence to be decoded. Accordingly, based on the generated output token(s), the output token sequencemay be produced.

226 204 228 230 202 202 226 228 202 202 In certain aspects, the output token(s), generated by token predictor(e.g., via language model), may be combined into a partial output token sequence(e.g., a version of the output token sequence) and used to update the input token sequencefor a next decoding cycle. For example, for an input token sequence“my favorite foods include,” token predictormay generate output tokens “spaghetti” and “and.” Output tokens “spaghetti” and “and” may be combined into a partial output token sequence“spaghetti and” and used to update input token sequencesuch that the updated input token sequenceincludes “my favorite foods include spaghetti and” for a next decoding cycle.

200 200 212 220 Systemachieves efficient and accurate token prediction by dynamically adapting to uncertainty trends and optimizing the balance between parallel and sequential decoding. Systemrealizes such advantages by comparing the entropy scoreof each token future position individually against the current adaptive entropy threshold. Future token positions that satisfy the threshold may be decoded in parallel, while those that do not may be decoded sequentially. Example Adaptive Entropy Threshold Scheduling and Token Prediction

3 FIG.A 2 FIG. 300 300 300 206 210 214 218 200 depicts an example processfor adaptive entropy threshold scheduling. In certain aspects, example processmay be used to select and/or adjust an adaptive entropy threshold for token prediction using set block decoding. In certain aspects, example processmay be performed by token generator, entropy scores generator, entropy monitor, and adaptive entropy schedulerof systemdepicted and described with respect to.

3 FIG.A 3 FIG.B 300 302 Although not meant to be limiting to this specific example, as shown inand, example processmay be used to select and/or adjust an adaptive entropy threshold for predicting token(s) based on an input token sequence“the tallest mountain above sea level is.”

302 204 302 304 1 304 2 304 3 304 4 304 304 308 1 308 2 308 3 308 4 308 304 302 3 FIG. 2 FIG. For example, input token sequence“the tallest mountain above sea level is” may be provided as input by a user. A language model (not shown in), such as language modelin, may process the input token sequenceand identify four future token positions-,-,-, and-(collectively referred to herein as “token positions”) as candidate positions for decoding. For each token position, the language model may generate a probability distribution-,-,-,-(collectively referred to herein as “probability distributions”), respectively, which reflects the likelihood of candidate output tokens appearing at respective token positions, such as based on the input token sequence.

308 1 304 1 308 2 304 2 308 3 304 3 308 4 304 4 304 1 304 1 302 304 2 304 2 302 304 1 304 3 304 3 302 304 1 304 2 304 4 304 4 302 304 1 304 2 304 3 For example, a probability distribution (PD)-may be generated for a vocabulary of candidate output tokens for token position-, a probability distribution-may be generated for a vocabulary of candidate output tokens for token position-, a probability distribution-may be generated for a vocabulary of candidate output tokens for token position-, and a probability distribution-may be generated for a vocabulary of candidate output tokens for token position-. For token position-, a candidate output token “Mount” may be assigned a highest probability score indicating that the candidate output token “Mount” is the most likely token for that token position-based on the input token sequence. For token position-, a candidate output token “Everest” may be assigned a highest probability score indicating that the candidate output token “Everest” is the most likely token for that token position-based on the input token sequenceand/or a token for token position-. For token position-, a candidate output token “in” may be assigned a highest probability score indicating that the candidate output token “in” is the most likely token for that token position-based on the input token sequence, a token for token position-, and/or a token for token position-. For token position-, a candidate output token “Asia” may be assigned a highest probability score indicating that the candidate output token “Asia” is the most likely token for that token position-based on the input token sequence, a token for token position-, a token for token position-, and/or a token for token position-.

t 310 310 308 310 1 308 1 310 2 308 2 310 3 308 3 310 4 308 4 308 310 308 310 308 1 310 1 308 3 310 3 Entropy scores (H)(also individually referred to herein as “entropy score”) may be generated from probability distributions. For example, an entropy score-may be generated based on probability distribution-, an entropy score-may be generated based on probability distribution-, an entropy score-may be generated based on probability distribution-, and an entropy score-may be generated based on probability distribution-. In certain aspects, a sharply peaked probability distributionmay yield a lower entropy score, while a flatter probability distributionmay yield a higher entropy score. For example, if probability distribution-places most of its confidence on “Mount,” entropy score-may be low. In another example, if probability distribution-spreads probability broadly across function words like “in,” “the,” and “is,” entropy score-may be relatively high.

310 304 310 1 310 4 304 1 304 4 312 312 310 1 310 4 304 1 304 4 Entropy scoresfor token positionsunder consideration (e.g., being evaluated) in the current decoding cycle, such as entropy scores-through-for token positions-through-, may be used to determine a current decoding cycle entropy score. In certain aspects, the current decoding cycle entropy scoremay represent a mean, a median, or a weighted aggregation of entropy scores-through-for token positions-through-.

314 312 314 3 FIG.A Trend metricsmay be computed through a series of actions designed to analyze uncertainty trends for the language model. First, a sliding window (not shown in) may be updated to include the current decoding cycle entropy scorealongside a set of previous decoding cycle entropy scores. In certain aspects, to reduce noise or jitter, the sliding window values may optionally undergo smoothing, using techniques such as an exponential moving average. Next, trend metricsmay be calculated based on the values in the sliding window. For example, in certain aspects, a slope of the decoding cycle entropy scores within the sliding window may be calculated, such as to capture the rate and/or direction of change in uncertainty. In certain aspects, variance of the decoding cycle entropy scores within the sliding window may be calculated, such as to measure the dispersion and/or spikiness in recent uncertainty levels of the language model.

316 302 314 314 316 316 314 316 314 316 316 316 A current adaptive entropy threshold(e.g., to use for set block decoding when performing token prediction based on input token sequence) may be determined based on trend metrics. For example, in certain aspects, a mapping between trend metricsand the current adaptive entropy thresholdmay be used to select the current adaptive entropy thresholdbased on trend metrics. In certain aspects, one or more rules may be applied to adjust a previous adaptive entropy threshold (e.g., for a previous decoding cycle) to obtain the current adaptive entropy threshold. In certain aspects, the rules may be applied based on whether the trend metricssatisfy specified criteria. For example, if the slope exceeds a threshold, indicating an entropy spike, a previous adaptive entropy threshold may be decreased to a current adaptive entropy threshold, such as to limit parallel decoding and improve accuracy. Conversely, if the slope is less than the threshold, the previous adaptive entropy threshold may be increased to a current adaptive entropy threshold, such as to enable greater parallelization. In certain aspects, variance may help to distinguish between a real trend (e.g., stable entropy but rising or falling) and noise (e.g., unstable entropy), thereby acting as a stabilizer as to how much an adaptive entropy threshold is increased and/or decreased. In certain aspects, decoding feedback, such as previous entropy scores, latency metrics, and/or accuracy metrics, may be used to determine a threshold adjustment to be applied to a previous adaptive entropy threshold (e.g., to increase or decrease the previous adaptive entropy threshold by an increment to obtain current adaptive entropy threshold). In certain aspects, bounds and/or hysteresis may be applied to prevent oscillation and constrain aggressiveness in threshold adjustments.

3 FIG.B 3 FIG.A 2 FIG. 3 FIG.A 350 350 222 226 200 302 depicts example token predictionusing the current adaptive entropy threshold determined in. In certain aspects, example token predictionmay be performed by entropy-based samplerand token predictorof system, depicted and described with respect to, such as to generate tokens for input token sequence“the tallest mountain above sea level is” (e.g., introduced in).

3 FIG.B 310 1 310 2 310 3 310 4 304 1 304 2 304 3 304 4 318 316 350 310 1 304 1 316 310 1 316 310 1 316 304 2 304 3 304 4 318 310 1 310 2 310 1 316 310 2 316 318 310 3 310 3 316 318 304 4 310 3 304 3 316 As shown in, entropy scores-,-,-,-for token positions-,-,-,-, respectively, may be compared, at, against current adaptive entropy threshold. For example, for example token prediction, entropy score-, generated for token position-, may be compared to current adaptive entropy thresholdto determine if entropy score-satisfies the current adaptive entropy threshold(e.g., determine if entropy score-< current adaptive entropy threshold), and thus qualifies for parallel decoding. Similar comparison may be performed for token positions-,-, and in some cases,-. In this example, at, it is determined that entropy score-and entropy score-satisfy the current adaptive entropy threshold (e.g., entropy score-< current adaptive entropy thresholdand entropy score-< current adaptive entropy threshold). Further, in this example, at, it is determined that entropy score-does not satisfy the current adaptive entropy threshold (e.g., entropy score-< current adaptive entropy threshold). In this example, the comparison (e.g., at) may not be performed for token position-due to the entropy score-for token position-not satisfying the current adaptive entropy threshold, such that only contiguous tokens are decoded in parallel.

304 1 304 2 310 1 310 2 316 304 1 304 2 304 1 304 2 Because token positions-and-have entropy scores-and-, respectively, that satisfy the current adaptive entropy threshold, set block decoding may be performed for token positions-and-to decode output tokens for token positions-and-in parallel.

304 3 304 4 304 3 304 4 304 3 304 4 Decoding output tokens for token positions-and-may occur in the same decoding cycle; however, tokens for token positions-and-may be decoded using single token decoding (instead of set block decoding). In other words, the remaining token positions in that same block (e.g., token positions-and-) may be decoded immediately afterward using next-token prediction within the same decoding cycle.

3 FIG.B 304 1 304 2 It is noted that althoughdepicts sequential left-to-right decoding (e.g., for set block decoding of tokens in token positions-and-), in some other examples, set block decoding may not require sequential left-to-right order decoding. In other words, any subset of token positions (e.g., in a block for a decoding cycle) may be decoded in parallel, and the remaining token positions may be decoded by next-token prediction or single token decoding in the same block/decoding cycle. When the entire block is exhausted, in a decoding cycle, the system may emit metric(s) to a reinforcement learning controller and begin the next decoding cycle.

304 1 304 2 302 320 300 350 304 3 304 3 304 1 304 2 304 3 304 4 152 150 1 FIG. As such, output tokens for token positions-and-may be generated as “Mount” and “Everet,” respectively. These output tokens and the input token sequencemay be combined into a first version of the output token sequencethat reads “the tallest mountain above sea level is Mount Everest.” In certain aspects, example processfor adaptive entropy threshold scheduling and example token predictionmay be repeated for one or more subsequent decoding cycles to decode tokens for the remaining token positions-and-. After tokens for all token positions-,-,-, and-have been generated, an end-of-sequence condition may be met, and a final output token sequence may be generated. In certain aspects, the final output token sequence may be generated for rendering on a UI of a computing device, such as UIof client devicedepicted and described with respect to.

4 FIG. 4 FIG. depicts example dynamic change in language model uncertainty over time, represented by entropy scores generated across decoding cycles. As shown in, initially, entropy remains low and stable, corresponding to predictable text segments such as common phrases or punctuation. As the sequence transitions into a more complex segment, such as a numeric computation or syntax-sensitive code snippet, entropy spikes sharply, indicating increased uncertainty. Once the challenging segment is resolved (e.g., decoded), entropy returns to a lower and steadier range again.

An adaptive entropy threshold may be dynamically adjusted based on these changes. For example, when the entropy curve is stable, with a flat or gently declining slope, the adaptive entropy threshold may be increased to enable larger or more frequent parallel blocks in decoding. Conversely, when the curve spikes, showing a positive slope and elevated variance, the adaptive entropy threshold may be decreased. This reduction may constrain parallelism, thereby helping to ensure accuracy, such as by focusing on sequential decoding for positions with higher uncertainty. This adaptive behavior may provide a balance between decoding efficiency and prediction accuracy.

5 FIG. 2 FIG. 7 FIG. 500 500 200 700 depicts an example methodfor token prediction. In one aspect, methodcan be implemented by the systemofand/or processing systemof.

500 505 Methodbegins at blockwith processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens.

500 510 Methodthen proceeds to blockwith determining, for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions.

500 515 Methodthen proceeds to blockwith performing, based on the determination for each respective token position of the set of token positions, a decoding operation comprising: a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence; or a single-token decoding operation to generate a single output token that corresponds to the first token position.

500 520 Methodthen proceeds to blockwith generating a first version of the output token sequence including the at least two output tokens or the single output token.

520 In some aspects, the set of token positions comprises at least two token positions, of the plurality of token positions, that correspond to at least the first token position and a second token position in the output token sequence, the decoding operation comprises the set block decoding operation based on the entropy score associated with each respective token position of at least two token positions satisfying the current adaptive entropy threshold, and blockincludes generating the first version of the output token sequence with the at least two output tokens.

In some aspects, performing the set block decoding operation comprises decoding, in parallel, the sets of candidate output tokens corresponding to the at least two token positions to generate the at least two output tokens.

520 In some aspects, the set of token positions comprises the first token position and the second token position in the output token sequence, the decoding operation comprises the single-token decoding operation based on: the entropy score associated with the first token position satisfying the current adaptive entropy threshold; and the entropy score associated with the second token position not satisfying the current adaptive entropy threshold, and blockincludes generating the first version of the output token sequence with the single output token.

520 In some aspects, the set of token positions comprises the first token position in the output token sequence, the decoding operation comprises the single-token decoding operation based on the entropy score associated with the first token position not satisfying the current adaptive entropy threshold, and blockincludes generating the first version of the output token sequence with the single output token.

500 In some aspects, methodfurther includes determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions.

500 In some aspects, methodfurther includes updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores.

500 In some aspects, methodfurther includes determining one or more trend metrics based on the sliding window of entropy scores.

500 In some aspects, methodfurther includes determining the current adaptive entropy threshold based on the trend metrics.

In some aspects, the current decoding cycle entropy score comprises: a mean entropy score; a median entropy score; or a weighted aggregation entropy score.

In some aspects, the one or more trend metrics comprise at least one of: a slope; or a variance.

In some aspects, determining the current adaptive entropy threshold comprises determining a first mapping between the one or more trend metrics and the current adaptive entropy threshold.

In some aspects, determining the current adaptive entropy threshold comprises determining one or more rules.

In some aspects, determining the current adaptive entropy threshold comprises: determining at least one of the one or more trend metrics satisfies at least one threshold; and adjusting, based on the determination, a previous adaptive entropy threshold by an increment.

In some aspects, the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is greater than a slope threshold, and adjusting the previous adaptive entropy threshold comprises decreasing the previous adaptive entropy threshold by the increment.

In some aspects, the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is less than a slope threshold, and adjusting the previous adaptive entropy threshold comprises increasing the previous adaptive entropy threshold by the increment.

500 In some aspects, the methodfurther comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and determining the current adaptive entropy threshold comprises adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback.

In some aspects, the decoding feedback comprises information about at least one of the plurality of previous decoding cycle entropy scores; a latency metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles; or an accuracy metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles.

500 500 500 By leveraging methodfor token prediction, significant technical advantages may be achieved. For example, methodprovides a systematic approach to dynamically compute and adjust an adaptive entropy threshold based on trend metrics derived from a sliding window of entropy scores across decoding cycles. This adaptive entropy threshold is uniformly applied to qualify token positions for parallel or single-token decoding. Adaptive gating enables increased parallelization in low-uncertainty regions, reducing latency and computational costs, while restricting parallel decoding in high-uncertainty regions to ensure accuracy and coherence. The approach minimizes forward passes during inference, mitigates error propagation by deferring ambiguous positions, and eliminates manual tuning through real-time threshold scheduling. When optional decoding feedback, such as latency and accuracy metrics, is incorporated, methodfurther refines threshold adjustments over time, enhancing throughput without compromising correctness. Accordingly, the performance, reliability, and efficiency of AI-driven token prediction workflows may be enhanced, supporting robust, scalable deployments and improved user experience.

5 FIG. Note thatis just one example of a method, and other methods including fewer, additional, or alternative operations are possible consistent with this disclosure.

6 FIG. 2 FIG. 8 FIG. 600 600 200 800 depicts an example methodfor adaptive entropy threshold adjustment for token prediction. In one aspect, methodcan be implemented by the systemofand/or processing systemof.

600 605 Methodbegins at blockwith processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens.

600 610 Methodthen proceeds to blockwith determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions.

600 615 Methodthen proceeds to blockwith updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores.

600 620 Methodthen proceeds to blockwith determining one or more trend metrics based on the sliding window of entropy scores.

600 625 Methodthen proceeds to blockwith determining a current adaptive entropy threshold based on the trend metrics.

600 630 Methodthen proceeds to blockwith predicting one or more tokens, associated with one or more token positions in the plurality of token positions, in the output token sequence based on the current adaptive entropy threshold.

In some aspects, the current decoding cycle entropy score comprises: a mean entropy score; a median entropy score; or a weighted aggregation entropy score.

In some aspects, the one or more trend metrics comprise at least one of: a slope; or a variance.

625 In some aspects, blockincludes determining a first mapping between the one or more trend metrics and the current adaptive entropy threshold.

625 In some aspects, blockincludes determining one or more rules.

625 In some aspects, blockincludes: determining at least one of the one or more trend metrics satisfies at least one threshold; and adjusting, based on the determination, a previous adaptive entropy threshold by an increment.

In some aspects, the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is greater than a slope threshold, and adjusting the previous adaptive entropy threshold comprises decreasing the previous adaptive entropy threshold by the increment.

In some aspects, the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is less than a slope threshold, and adjusting the previous adaptive entropy threshold comprises increasing the previous adaptive entropy threshold by the increment.

600 625 In some aspects, the methodfurther comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and blockincludes adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback.

In some aspects, the decoding feedback comprises information about at least one of: the plurality of previous decoding cycle entropy scores; a latency metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles; or an accuracy metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles.

600 600 600 By leveraging methodfor adaptive entropy threshold adjustment, significant technical advantages may be achieved. For example, methodenables dynamic computation and adjustment of an adaptive entropy threshold based on trend metrics derived from a sliding window of entropy scores across decoding cycles. This threshold is applied to qualify token positions for prediction, allowing for parallel decoding in low-uncertainty regions to reduce latency and computational costs, while restricting decoding in high-uncertainty regions to ensure accuracy and coherence. The approach minimizes error propagation, eliminates manual tuning through real-time threshold scheduling, and, in some cases, incorporates decoding feedback, such as latency and accuracy metrics, to refine adjustments over time. Accordingly, methodenhances the performance, reliability, and efficiency of token prediction workflows, supporting robust and scalable deployments across diverse applications.

6 FIG. Note thatis just one example of a method, and other methods including fewer, additional, or alternative operations are possible consistent with this disclosure.

7 FIG. 5 FIG. 700 500 depicts an example processing systemconfigured to perform various aspects described herein, including, for example, methodas described above with respect to.

700 Processing systemis an example of an electronic device configured to execute computer-executable instructions, such as those derived from compiled computer code, including without limitation personal computers, tablet computers, servers, smart phones, smart devices, wearable devices, augmented and/or virtual reality devices, and others.

700 702 704 706 708 700 712 710 710 In the depicted example, processing systemincludes one or more processors, one or more input/output devices, one or more display devices, one or more network interfacesthrough which processing systemis connected to one or more networks (e.g., a local network, an intranet, the Internet, or any other group of processing systems communicatively connected to each other), and computer-readable medium. In the depicted example, the aforementioned components are coupled by a bus, which may generally be configured for data exchange amongst the components. Busmay be representative of multiple buses, while only one is depicted for simplicity.

702 712 702 712 710 702 706 708 712 702 Processor(s)are generally configured to retrieve and execute instructions stored in one or more memories, including local memories like computer-readable medium, as well as remote memories and data stores. Similarly, processor(s)are configured to store application data residing in local memories like the computer-readable medium, as well as remote memories and data stores. More generally, busis configured to transmit programming instructions and application data among the processor(s), display device(s), network interface(s), and/or computer-readable medium. In certain embodiments, processor(s)are representative of one or more central processing units (CPUs), graphics processing unit (GPUs), tensor processing unit (TPUs), accelerators, and other processing devices.

704 700 700 704 Input/output device(s)may include any device, mechanism, system, interactive display, and/or various other hardware and software components for communicating information between processing systemand a user of processing system. For example, input/output device(s)may include input hardware, such as a keyboard, touch screen, button, microphone, speaker, and/or other device for receiving inputs from the user and sending outputs to the user.

706 706 706 706 Display device(s)may generally include any sort of device configured to display data, information, graphics, user interface elements, and the like to a user. For example, display device(s)may include internal and external displays such as an internal display of a tablet computer or an external display for a server computer or a projector. Display device(s)may further include displays for devices, such as augmented, virtual, and/or extended reality devices. In various embodiments, display device(s)may be configured to display a graphical user interface.

708 700 708 708 Network interface(s)provide processing systemwith access to external networks and thereby to external processing systems. Network interface(s)can generally be any hardware and/or software capable of transmitting and/or receiving data via a wired or wireless network connection. Accordingly, network interface(s)can include a communication transceiver for sending and/or receiving any wired and/or wireless communication.

712 712 714 716 718 720 722 724 726 728 730 714 730 700 500 5 FIG. Computer-readable mediummay be a volatile memory, such as a random access memory (RAM), or a nonvolatile memory, such as nonvolatile random access memory (NVRAM), or the like. In this example, computer-readable mediumincludes processing component, determining component, performing component, generating component, decoding component, updating component, adjusting component, decreasing component, increasing component. Processing of the components-may enable and cause the processing systemto perform the methoddescribed with respect to, or any aspect related to it.

714 505 716 510 718 515 720 520 5 FIG. 5 FIG. 5 FIG. 5 FIG. In certain embodiments, processing componentis configured to process an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens, as described inwith reference to block. In certain embodiments, determining componentis configured to determine, for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions, as described inwith reference to block. In certain embodiments, performing componentis configured to perform, based on the determination for each respective token position of the set of token positions, a decoding operation comprising: a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence; or a single-token decoding operation to generate a single output token that corresponds to the first token position, as described inwith reference to block. In certain embodiments, generating componentis configured to generate a first version of the output token sequence including the at least two output tokens or the single output token, as described inwith reference to block.

7 FIG. Note thatis just one example of a processing system consistent with aspects described herein, and other processing systems having additional, alternative, or fewer components are possible consistent with this disclosure.

8 FIG. 6 FIG. 800 600 depicts an example processing systemconfigured to perform various aspects described herein, including, for example, methodas described above with respect to.

800 Processing systemis an example of an electronic device configured to execute computer-executable instructions, such as those derived from compiled computer code, including without limitation personal computers, tablet computers, servers, smart phones, smart devices, wearable devices, augmented and/or virtual reality devices, and others.

800 802 804 806 808 800 812 810 810 In the depicted example, processing systemincludes one or more processors, one or more input/output devices, one or more display devices, one or more network interfacesthrough which processing systemis connected to one or more networks (e.g., a local network, an intranet, the Internet, or any other group of processing systems communicatively connected to each other), and computer-readable medium. In the depicted example, the aforementioned components are coupled by a bus, which may generally be configured for data exchange amongst the components. Busmay be representative of multiple buses, while only one is depicted for simplicity.

802 812 802 812 810 802 806 808 812 802 Processor(s)are generally configured to retrieve and execute instructions stored in one or more memories, including local memories like computer-readable medium, as well as remote memories and data stores. Similarly, processor(s)are configured to store application data residing in local memories like the computer-readable medium, as well as remote memories and data stores. More generally, busis configured to transmit programming instructions and application data among the processor(s), display device(s), network interface(s), and/or computer-readable medium. In certain embodiments, processor(s)are representative of one or more central processing units (CPUs), graphics processing unit (GPUs), tensor processing unit (TPUs), accelerators, and other processing devices.

804 800 800 804 Input/output device(s)may include any device, mechanism, system, interactive display, and/or various other hardware and software components for communicating information between processing systemand a user of processing system. For example, input/output device(s)may include input hardware, such as a keyboard, touch screen, button, microphone, speaker, and/or other device for receiving inputs from the user and sending outputs to the user.

806 806 806 806 Display device(s)may generally include any sort of device configured to display data, information, graphics, user interface elements, and the like to a user. For example, display device(s)may include internal and external displays such as an internal display of a tablet computer or an external display for a server computer or a projector. Display device(s)may further include displays for devices, such as augmented, virtual, and/or extended reality devices. In various embodiments, display device(s)may be configured to display a graphical user interface.

808 800 808 808 Network interface(s)provide processing systemwith access to external networks and thereby to external processing systems. Network interface(s)can generally be any hardware and/or software capable of transmitting and/or receiving data via a wired or wireless network connection. Accordingly, network interface(s)can include a communication transceiver for sending and/or receiving any wired and/or wireless communication.

812 812 814 816 818 820 822 824 826 828 814 828 800 600 6 FIG. Computer-readable mediummay be a volatile memory, such as a random access memory (RAM), or a nonvolatile memory, such as nonvolatile random access memory (NVRAM), or the like. In this example, computer-readable mediumincludes processing component, determining component, updating component, predicting component, adjusting component, decreasing component, increasing component, and obtaining component. Processing of the components-may enable and cause the processing systemto perform the methoddescribed with respect to, or any aspect related to it.

814 605 816 610 818 615 816 620 816 625 820 630 6 FIG. 6 FIG. 6 FIG. 6 FIG. 6 FIG. 6 FIG. In certain embodiments, processing componentis configured to process an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens, as described inwith reference to block. In certain embodiments, determining componentis configured to determine a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions, as described inwith reference to block. In certain embodiments, updating componentis configured to update a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores, as described inwith reference to block. In certain embodiments, determining componentis configured to determine one or more trend metrics based on the sliding window of entropy scores, as described inwith reference to block. In certain embodiments, determining componentis configured to determine a current adaptive entropy threshold based on the trend metrics, as described inwith reference to block. In certain embodiments, predicting componentis configured to predict one or more tokens, associated with one or more token positions in the plurality of token positions, in the output token sequence based on the current adaptive entropy threshold, as described inwith reference to block.

8 FIG. Note thatis just one example of a processing system consistent with aspects described herein, and other processing systems having additional, alternative, or fewer components are possible consistent with this disclosure.

Implementation examples are described in the following numbered clauses:

Clause 1: A method of token prediction, comprising: processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each respective candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens; determining, for each respective token position of a set of token positions from the plurality of token positions in the output token sequence, whether the entropy score associated with the respective token position satisfies a current adaptive entropy threshold, wherein the current adaptive entropy threshold is based on the entropy score associated with each token position of the plurality of token positions; performing, based on the determination for each respective token position of the set of token positions, a decoding operation comprising: a set block decoding operation to generate at least two output tokens that correspond to at least a first token position and a second token position of the plurality of token positions in the output token sequence; or a single-token decoding operation to generate a single output token that corresponds to the first token position; and generating a first version of the output token sequence including the at least two output tokens or the single output token.

Clause 2: The method of Clause 1, wherein: the set of token positions comprises at least two token positions, of the plurality of token positions, that correspond to at least the first token position and a second token position in the output token sequence, the decoding operation comprises the set block decoding operation based on the entropy score associated with each respective token position of at least two token positions satisfying the current adaptive entropy threshold, and generating the first version of the output token sequence comprises generating the first version of the output token sequence with the at least two output tokens.

Clause 3: The method of Clause 2, wherein performing the set block decoding operation comprises decoding, in parallel, the sets of candidate output tokens corresponding to the at least two token positions to generate the at least two output tokens.

Clause 4: The method of any one of Clauses 1-3, wherein: the set of token positions comprises the first token position and the second token position in the output token sequence, the decoding operation comprises the single-token decoding operation based on: the entropy score associated with the first token position satisfying the current adaptive entropy threshold; and the entropy score associated with the second token position not satisfying the current adaptive entropy threshold, and generating the first version of the output token sequence comprises generating the first version of the output token sequence with the single output token.

Clause 5: The method of any one of Clauses 1-4, wherein: the set of token positions comprises the first token position in the output token sequence, the decoding operation comprises the single-token decoding operation based on the entropy score associated with the first token position not satisfying the current adaptive entropy threshold, and generating the first version of the output token sequence comprises generating the first version of the output token sequence with the single output token.

Clause 6: The method of any one of Clauses 1-5, further comprising: determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions; updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores; determining one or more trend metrics based on the sliding window of entropy scores; and determining the current adaptive entropy threshold based on the trend metrics.

Clause 7: The method of Clause 6, wherein the current decoding cycle entropy score comprises: a mean entropy score; a median entropy score; or a weighted aggregation entropy score.

Clause 8: The method of Clause 6, wherein the one or more trend metrics comprise at least one of: a slope; or a variance.

Clause 9: The method of Clause 6, wherein determining the current adaptive entropy threshold comprises determining a first mapping between the one or more trend metrics and the current adaptive entropy threshold.

Clause 10: The method of Clause 6, wherein determining the current adaptive entropy threshold comprises determining one or more rules.

Clause 11: The method of Clause 6, wherein determining the current adaptive entropy threshold comprises: determining at least one of the one or more trend metrics satisfies at least one threshold; and adjusting, based on the determination, a previous adaptive entropy threshold by an increment.

Clause 12: The method of Clause 11, wherein: the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is greater than a slope threshold, and adjusting the previous adaptive entropy threshold comprises decreasing the previous adaptive entropy threshold by the increment.

Clause 13: The method of Clause 11, wherein: the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is less than a slope threshold, and adjusting the previous adaptive entropy threshold comprises increasing the previous adaptive entropy threshold by the increment.

Clause 14: The method of Clause 6, wherein: the method further comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and determining the current adaptive entropy threshold comprises adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback.

Clause 15: The method of Clause 14, wherein the decoding feedback comprises information about at least one of: the plurality of previous decoding cycle entropy scores; a latency metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles; or an accuracy metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles.

Clause 16: A method of adaptive entropy threshold adjustment for token prediction, comprising: processing an input token sequence to generate, for each respective token position of a plurality of token positions in an output token sequence: a probability distribution for a set of candidate output tokens associated with the respective token position, wherein the probability distribution comprises a probability score for each candidate output token in the set of candidate output tokens; and an entropy score based on the probability distribution for the set of candidate output tokens; determining a current decoding cycle entropy score based on the entropy score associated with each token position of the plurality of token positions; updating a sliding window of entropy scores to include the current decoding cycle entropy score, the sliding window of entropy scores comprising a plurality of previous decoding cycle entropy scores; determining one or more trend metrics based on the sliding window of entropy scores; determining a current adaptive entropy threshold based on the trend metrics; and predicting one or more tokens, associated with one or more token positions in the plurality of token positions, in the output token sequence based on the current adaptive entropy threshold.

Clause 17: The method of Clause 16, wherein the current decoding cycle entropy score comprises: a mean entropy score; a median entropy score; or a weighted aggregation entropy score.

Clause 18: The method of any one of Clauses 16-17, wherein the one or more trend metrics comprise at least one of: a slope; or a variance.

Clause 19: The method of any one of Clauses 16-18, wherein determining the current adaptive entropy threshold comprises determining a first mapping between the one or more trend metrics and the current adaptive entropy threshold.

Clause 20: The method of any one of Clauses 16-19, wherein determining the current adaptive entropy threshold comprises determining one or more rules.

Clause 21: The method of any one of Clauses 16-20, wherein determining the current adaptive entropy threshold comprises: determining at least one of the one or more trend metrics satisfies at least one threshold; and adjusting, based on the determination, a previous adaptive entropy threshold by an increment.

Clause 22: The method of Clause 21, wherein: the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is greater than a slope threshold, and adjusting the previous adaptive entropy threshold comprises decreasing the previous adaptive entropy threshold by the increment.

Clause 23: The method of Clause 21, wherein: the one or more trend metrics comprise a slope, determining the at least one of the one or more trend metrics satisfies the at least one threshold comprises determining the slope is less than a slope threshold, and adjusting the previous adaptive entropy threshold comprises increasing the previous adaptive entropy threshold by the increment.

Clause 24: The method of any one of Clauses 16-23, wherein: the method further comprises obtaining decoding feedback for a plurality of previous decoding cycles associated with the plurality of previous decoding cycle entropy scores, and determining the current adaptive entropy threshold comprises adjusting a previous adaptive entropy threshold by an increment based on the decoding feedback.

Clause 25: The method of Clause 24, wherein the decoding feedback comprises information about at least one of: the plurality of previous decoding cycle entropy scores; a latency metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles; or an accuracy metric associated with each respective previous decoding cycle of the plurality of previous decoding cycles.

Clause 26: A processing system, comprising: memory comprising computer-executable instructions; and one or more processors configured to execute the computer-executable instructions and cause the processing system to perform a method in accordance with any one of Clauses 1-25.

Clause 27: A processing system, comprising means for performing a method in accordance with any one of Clauses 1-25.

Clause 28: A non-transitory computer-readable medium storing program code for causing a processing system to perform the steps of any one of Clauses 1-25.

Clause 29: A computer program product embodied on a computer-readable storage medium comprising code for performing a method in accordance with any one of Clauses 1-25.

The preceding description is provided to enable any person skilled in the art to practice the various embodiments described herein. The examples discussed herein are not limiting of the scope, applicability, or embodiments set forth in the claims. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments. For example, changes may be made in the function and arrangement of elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components as appropriate. For instance, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects set forth herein. In addition, the scope of the disclosure is intended to cover such an apparatus or method that is practiced using other structure, functionality, or structure and functionality in addition to, or other than, the various aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

As used herein, a phrase referring to “at least one of” a list of items refers to any combination of those items, including single members. As an example, “at least one of: a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c, as well as any combination with multiples of the same element (e.g., a-a, a-a-a, a-a-b, a-a-c, a-b-b, a-c-c, b-b, b-b-b, b-b-c, c-c, and c-c-c or any other ordering of a, b, and c).

As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure), ascertaining and the like. Also, “determining” may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and the like. Also, “determining” may include resolving, selecting, choosing, establishing and the like.

The methods disclosed herein comprise one or more steps or actions for achieving the methods. The method steps and/or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and/or use of specific steps and/or actions may be modified without departing from the scope of the claims. Further, the various operations of methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and/or software component(s) and/or module(s), including, but not limited to a circuit, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

The following claims are not intended to be limited to the embodiments shown herein, but are to be accorded the full scope consistent with the language of the claims. Within a claim, reference to an element in the singular is not intended to mean “one and only one” unless specifically so stated, but rather “one or more.” Unless specifically stated otherwise, the term “some” refers to one or more. No claim element is to be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase “means for” or, in the case of a method claim, the element is recited using the phrase “step for.” All structural and functional equivalents to the elements of the various aspects described throughout this disclosure that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether such disclosure is explicitly recited in the claims.

Classification Codes (CPC)

Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.

Patent Metadata

Filing Date

December 17, 2025

Publication Date

July 28, 2026

Inventors

Sagiv Antebi
Matan Vetzler
Shai Ardazi
Ofir Ben Shoham

Want to explore more patents?

Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.

Citation & reuse

Analysis on this page is generated by Patentable — an AI-powered patent intelligence platform. AI-generated summaries, explanations, and analysis may be reused with attribution and a visible link back to the canonical URL below. Patent abstracts and claims are USPTO public domain.

Cite as: Patentable. “Entropy-based set block decoding utilizing an adaptive entropy threshold” (US-12694213-B2). https://patentable.app/patents/US-12694213-B2

© 2026 Patentable. All rights reserved.

Patentable is a research and drafting-assistant tool, not a law firm, and does not provide legal advice. Documents we generate are drafts for review by a licensed patent attorney.

Entropy-based set block decoding utilizing an adaptive entropy threshold — Sagiv Antebi | Patentable