A method for quantizing a machine learning model during an inference phase, including determining a normalization factor using a set of floating-point values and a damped value of a damped value sequence; and assigning a quantized value for each floating-point value of the set of floating-point values based on the damped value sequence and the normalization factor.
Legal claims defining the scope of protection, as filed with the USPTO.
2. The method of claim 1, further comprising determining the damped value from the damped value sequence using the damped value sequence and a number of quantization bits.
3. The method of claim 2, wherein determining the damped value from the damped value sequence comprises using the largest damped value from the damped value sequence based on the number of quantization bits.
6. The method of claim 1, wherein the normalization factor is determined using a maximum value from the set of floating-point values, a minimum value from the set of floating-point values, and the damped value.
11. The computer-readable medium according to claim 10, wherein determining the damped value from the damped value sequence comprises using the largest damped value from the damped value sequence based on the number of quantization bits.
14. The computer-readable medium according to claim 9, wherein the normalization factor is determined using a maximum value from the set of floating-point values, a minimum value from the set of floating-point values, and the damped value.
18. The quantizer of claim 17, wherein the damped value from the damped value sequence is determined using the damped value sequence and a number of quantization bits.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
November 7, 2018
October 18, 2022
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.