Generally discussed herein are devices, systems, and methods for machine learning (ML) modeling of a system that operates on a multivector object. A method includes receiving, by an ML model, the multivector object as an input that represents a state of the multivector system. The method includes operating, by the ML model and using a Clifford layer that includes neurons that implement a multivector kernel, on the multivector input to generate a multivector output that represents the state of the multivector system responsive to the multivector input.
Legal claims defining the scope of protection, as filed with the USPTO.
receiving the multivector object that represents a state of the multivector system; and transforming, using a Clifford layer of the ML model that includes neurons that implement a multivector kernel, the multivector object to a multivector output that represents the state of the multivector system responsive to the multivector object, the Clifford layer is: (i) a Clifford Fourier Neural Operator (FNO) that includes a Clifford Fourier layer and a first Clifford convolution layer; or (ii) a second Clifford convolution layer wherein a kernel for the second Clifford convolution layer is constrained such that the second Clifford convolution layer is equivariant. . A method for machine learning (ML) modeling of a multivector system that operates on a multivector object, the method performed by an ML model, the method comprising:
claim 1 . The method of, wherein the Clifford layer is the Clifford FNO that includes the Clifford Fourier layer and the first Clifford convolution layer.
claim 1 . The method of, wherein the multivector system is an electromagnetic field or a dynamic fluid.
claim 1 . The method of, wherein the multivector object includes two or more of a scalar, a vector, a bi-vector, or a tri-vector.
claim 4 . The method of, wherein the multivector object includes the vector.
claim 4 . The method of, wherein the multivector object includes the bi-vector.
claim 4 . The method of, wherein the multivector object includes the tri-vector.
claim 1 . The method of, wherein the Clifford layer is the second Clifford convolution layer and the method further comprises constraining the kernel for the second Clifford convolution layer such that the second Clifford convolution layer is equivariant.
processing circuitry; and one or more memories including parameters for neurons trained for machine learning (ML) modeling of a multivector system that operates on a multivector object and instructions that, when executed by the processing circuitry, cause the processing circuitry to perform operations for transforming, by an ML model, a multivector object, the operations comprising: receiving the multivector object that represents a state of the multivector system; and transforming, by a Clifford layer of the ML model that includes neurons that implement a multivector kernel, the multivector object to a multivector output that represents the state of the multivector system responsive to the multivector object, the Clifford layer is: (i) a Clifford Fourier Neural Operator (FNO) that includes a Clifford Fourier layer and a first Clifford convolution layer; or (ii) a second Clifford convolution layer wherein a kernel for the second Clifford convolution layer is constrained such that the second Clifford convolution layer is equivariant. . A system comprising:
claim 9 . The system of, wherein the Clifford layer is the Clifford FNO that includes the Clifford Fourier layer and the first Clifford convolution layer.
claim 9 . The system of, wherein the multivector system is an electromagnetic field or a dynamic fluid.
claim 9 . The system of, wherein the multivector object includes two or more of a scalar, a vector, a bi-vector, or a tri-vector.
claim 9 . The system of, wherein the Clifford layer is the second Clifford convolution layer and the operations further include constraining the kernel for the second Clifford convolution layer such that the second Clifford convolution layer is equivariant.
receiving the multivector object that represents a state of the multivector system; and operating, by a Clifford layer of the ML model that includes neurons that implement a multivector kernel, on the multivector object to generate a multivector output that represents the state of the multivector system responsive to the multivector object, the Clifford layer is: (i) a Clifford Fourier Neural Operator (FNO) that includes a Clifford Fourier layer and a first Clifford convolution layer; or (ii) a second Clifford convolution layer wherein a kernel for the second Clifford convolution layer is constrained such that the second Clifford convolution layer is equivariant. . A computer-readable medium including instructions that, when executed by a machine, cause the machine to perform operations for implementing a machine learning (ML) model that models a multivector system that operates on a multivector object, the operations comprising:
claim 14 . The computer-readable medium of, wherein the Clifford layer is the Clifford FNO that includes Clifford Fourier layer and the first Clifford convolution layer.
claim 14 . The computer-readable medium of, wherein the multivector system is an electromagnetic field or a dynamic fluid.
claim 14 . The computer-readable medium of, wherein the multivector object includes two or more of a scalar, a vector, a bi-vector, or a tri-vector.
Complete technical specification and implementation details from the patent document.
This application claims the benefit of priority under 35 U.S.C. § 119(e), to U.S. Provisional Patent Application Ser. No. 63/404,810, titled “CLIFFORD NEURAL LAYERS FOR MULTI-VECTOR SYSTEM MODEL” filed on Sep. 8, 2022, which is hereby incorporated by reference in its entirety.
Current machine learning (ML) models operate on a standard vector or scalar
input. A vector includes multiple scalar entries.
A device, system, method, and computer-readable medium configured for implementing an ML model that operates on a multivector input are provided. The ML model can implement a multivector valued neural network (NN) which operates on multivector inputs. Clifford algebras are used to define how to operate on and between multivector objects. The output can be a multivector, classification, vector, or the like. One or more of a Clifford convolution layer, a Clifford Fourier transform layer, and an equivariant Clifford layer can be used to operate on the multivector input. Examples of systems that are modeled using multivectors include mechanical engineering, civil engineering, geophysics, or meteorology that use fluid dynamics (e.g., two-dimensional (2D) Navier-Stokes equations), electric, optical, or radio technologies that rely on solving three-dimensional (3D) Maxwell equations, and the modeling of other physical systems in general.
In the following description, reference is made to the accompanying drawings that form a part hereof, and in which is shown by way of illustration specific embodiments which may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It is to be understood that other embodiments may be utilized and that structural, logical, and/or electrical changes may be made without departing from the scope of the embodiments. The following description of embodiments is, therefore, not to be taken in a limited sense, and the scope of the embodiments is defined by the appended claims.
Some systems are classically modeled using standard vectors and scalars. These systems can benefit from modeling using multivector objects. The elements of a Clifford algebra are called multivectors or multivector objects. Multivector objects contain elements of subspaces (e.g., scalars, vectors, bivectors, trivectors, etc.). The geometric product is an operation which combines multivector objects such that the result is another multivector object. Embodiments can use multivector kernels to define convolutions over multivector object inputs. Embodiments can use the geometric product in the Fourier space to define Clifford Fourier layers over multivector objects.
The values of a given element of a multivector object are dependent on one or more values of other elements of the multivector object. No single, current ML model accounts for such diverse input elements or dependencies between the elements. Embodiments operate with neural network (NN) elements that operate using a Clifford algebra. This is in contrast to normal NN elements that operate on wx+b using an activation function, where w is a learned weight, x is a vector or scalar input value (but not a combination thereof), and b is a bias.
An application of Clifford neural layers is modeling through solving partial differential equations (PDEs). PDEs have significant use in science and engineering. PDEs simulate physical processes as scalar and vector fields interacting and coevolving over time. Due to the computationally expensive nature of standard PDE solution methods, neural PDE surrogates have become an active research topic to accelerate these simulations. However, current methods do not explicitly take into account the relationship between different fields and their internal components, which are often correlated. Modeling the time evolution of such correlated fields using multivector fields accounts for the relationship between different fields and their internal components. Multivector fields can include scalar, vector, as well as higher-order components, such as bivectors and trivectors. Algebraic properties of multivector fields, such as multiplication, addition, and other arithmetic operations can be described by Clifford algebras. Embodiments present the first usage of such multivector representations together with Clifford convolutions and Clifford Fourier transforms in the context of deep learning (DL). Resulting Clifford neural layers are universally applicable and can be used in the areas of fluid dynamics, weather forecasting, and the modeling of physical systems in general. Empirical evaluation of the benefit of Clifford neural layers is explored by replacing convolution and Fourier operations in common neural PDE surrogates with their Clifford counterparts. The PDE surrogates analyzed include two-dimensional Navier-Stokes and weather modeling tasks, as well as three-dimensional Maxwell equations. Clifford neural layers consistently improve generalization capabilities of the tested neural PDE surrogates.
Most scientific phenomena are described by the evolution and interaction of physical quantities over space and time. The concept of fields is one widely used construct to continuously parameterize these quantities over chosen coordinates. Prominent examples include (i) fluid mechanics, which has applications in domains ranging from mechanical and civil engineering, to geophysics and meteorology, and (ii) electromagnetism, which provides mathematical models for electric, optical, or radio technologies. The underlying equations of these examples are described in various forms of the Navier-Stokes equations and Maxwell's equations. For the majority of these equations, solutions are analytically intractable, and obtaining accurate predictions necessitates falling back on numerical approximation schemes often with prohibitive computation costs. DL's success at overcoming the curse of dimensionality in many fields has led to a surge of interest in scientific applications, especially at augmenting and replacing numerical solving schemes in fluid dynamics with neural networks.
Weather simulations are used to ground further description. In weather simulations, two different kinds of fields emerge: scalar fields such as temperature or humidity, and vector fields such as wind velocity or pressure gradients. Current DL-based approaches treat different vector field components the same as scalar fields, and stack all scalar fields along the channel dimension, thereby omitting the geometric relations between different components, both within vector fields as well as between individual vector and scalar fields. This practice leaves out important inductive bias information present in the input data. For example, wind velocities in the x- and y-directions are strongly related, such that they form a vector field. Additionally, the wind vector field and the scalar pressure field are related since the gradient of the pressure field causes air movement and subsequently influences the wind components. In this work, neural PDE surrogates are built which not only model local and global relationships, as current methods do, but also the relation between different fields (e.g., wind and pressure field) and field components (e.g. x- and y-component of the wind velocities).
1 2 3 e, e, e Basis vectors of the generating vector space of the Clifford algebra. i j e∧e i j Wedge (outer) product of basis vectors eand e. i j 1 j e· e= <e, e> i j Inner product of basis vectors eand e. 1 2 3 1 ee, ee, Basis bivectors of the vector space of the Clifford 2 3 ee algebra. 1 2 3 eee Basis trivector of the vector space of the Clifford algebra. 2 1 2 i= ee Pseudoscalar for Clifford algebras of grade 2. 3 1 2 3 i= eee Pseudoscalar for Clifford algebras of grade 3. x n Euclidean vector in R. x∧y wedge (outer) product of Euclidean vectors x and y. x · y = <x, y> Inner product of vectors x and y. a Multivector. ab Geometric product of multivectors a and b. î, ĵ, {circumflex over (k)} Base elements of quaternions.
1 FIG. 100 102 100 104 102 104 102 108 106 106 104 102 illustrates, by way of example, a diagram of an embodiment of an ML systemfor operating on a multivector object input. The ML systemas illustrated includes an ML modelthat receives and operates on the multivector object input. The ML modeloperates on the multivector object inputusing one or more Clifford neural layersto generate an output. The outputis a representation of a state of a system, being modeled by the ML model, responsive to a stimulus represented by the multivector object input.
102 102 The multivector object inputincludes a combination of entries comprising two or more of a scalar, vector, bi-vector, tri-vector, or other pseudo scalar. That is, the entries of the multivector object inputinclude two or more of a scalar, vector, bi-vector, tri-vector, or other higher-dimensional vector.
104 102 104 108 102 108 108 102 106 108 The ML modelmodels a system, such as a physical system, that includes a stimulus represented by the multivector object input. Examples of such physical systems includes that include 3D electromagnetic fields, dynamic fluids, and weather forecasting, among many others. The ML modelincludes the Clifford neural layerthat operates on the input. The Clifford neural layercan include a Clifford convolution layer, a Clifford Fourier layer, a Clifford equivariant layer, or a combination thereof, such as a Clifford Fourier Neural Operator (FNO) that includes a Clifford convolution layer and a Clifford Fourier layer. The Clifford neural layeroperates using a Clifford algebra on the multivector input. The outputis based on a result of the operations performed by the Clifford neural layer. More details on each of these layers can be found in the Appendix.
106 106 104 102 The outputcan include a scalar, classification, a multivector object, a vector, a combination thereof, or the like. The output, in general, represents the state of the system being modeled by the ML modelresponsive to the input.
Clifford algebras are at the core intersection of geometry and algebra, introduced to simplify spatial and geometrical relations between many mathematical concepts. For example, Clifford algebras naturally unify real numbers, vectors, complex numbers, quaternions, exterior algebras, and many more. Most notably, in contrast to standard vector analysis where primitives are scalars and vectors, Clifford algebras have additional spatial primitives for representing plane and volume segments. An expository example is the cross-product of two vectors in 3 dimensions, which naturally translates to a plane segment spanned by these two vectors. The cross product is often represented as a vector because it has three independent components but has a sign flip under reflection that a true vector does not. In Clifford algebras, different spatial primitives can be summarized into objects called multivectors.
2 FIG. 2 FIG. illustrates, by way of example, a diagram of the spatial primitives of the multivector object and an example multivector object. As illustrated in, the spatial primitives include a scalar, which is a one-dimensional primitive, a vector, which is a point with a direction, a bivector, which is a combination of two vectors, and a trivector, which is a combination of three vectors. There are higher-dimensional primitives. The vector is a multivector of grade 1, the bivector is a multivector of grade 2, etc. and there are multivectors of grade 4, and so on. The multivector includes respective entries corresponding to each of the scalar, vector, bivector, and trivector. Note that a given entry can be zero if the corresponding primitive is not present. The multivector
In embodiments, operations over feature fields in DL architectures are replaced by their Clifford algebra counterparts, which operate on multivector feature fields. For example, a convolutional kernel is endowed with multivector components, such that the kernel can convolve over multivector feature maps. Introducing Clifford neural layers yields at least two advantages: (i) scalar and vector components can be naturally grouped together into a single multivector object, with operations on, and mappings between multivectors defined by Clifford algebras; and (ii) higher order geometric information such as bivectors and trivectors arise, which match aforementioned higher order objects. For example, the curl of a vector field can be described via bivectors, and thus it is beneficial to have bivector filter components to match those.
Embodiments, more specifically, define Clifford convolution and Clifford Fourier transforms over multivector fields, building on connections between Clifford algebras, complex numbers, and quaternions. Based on these definitions, Clifford layers were implemented for different Clifford algebras, replacing convolution and Fourier operations in DL architectures. An evaluation of the effect of using Clifford based architectures on 2-dimensional problems of fluid mechanics and weather modeling, as well as on the 3-dimensional Maxwell's equations is provided to demonstrate that Clifford neural layers consistently improve generalization capabilities of the tested architectures.
3 5 FIGS.- 3 FIG. 3 FIG. 300 300 illustrate, by way of example, respective vector fields that help explain differences between prior treatment of fields and treatment of fields using multivector objects.illustrates a first vector fieldthat represents some physical phenomena. The vector fieldincludes a plurality of sinks and sources. A sink is a point to which vectors tend to settle (a point towards which vectors point). A sink is considered a point of stability. A source is a point from which vectors tend to point away or “originate” from. A source is considered a point of instability. Inthere is a batch size of 1, there are two channels (x and y), and there are 32 samples of each of x and y in the batch.
4 FIG. 4 FIG. 400 300 300 300 400 illustrates, by way of example, a vector fieldgenerated by decomposing the vector fieldinto respective x and y Fourier transforms and then attempting to reconstruct the vector fieldfrom the Fourier transforms. As can be seen, most, if not all, of the point source and sink structure from the vector fieldhas been lost in the vector field. Inthere is a batch size of 1, there are two channels (x and y), and there are 32 samples of each of x and y in the batch.
5 FIG. 5 FIG. 500 300 300 300 500 illustrates, by way of example, a vector fieldgenerated by treating the vectors fieldas a multivector object, performing a Clifford Fourier transform, using a Clifford Fourier layer, on the multivector object, and then reconstructing the vector fieldfrom the Clifford Fourier transform. As can be seen, most, if not all, of the point source and sink structure from the vector fieldhas been retained in the vector field. Inthere is a batch size of 1, there is a single, complex valued channel, and there are 32 samples of each of the single channel in the batch. The Fourier transform is taken over the single channel instead of multiple channels as is typically done.
The channel reduction is a way that the Clifford neural layers of embodiments provide improvements. The channel reduction simplifies the processing and retains more structure than breaking the problems into more channels, operating on the channels individually, and then recombining the outputs of the operations on the channels individually.
400 500 A reason the vector fielddid not retain the structure is because the x and y components were treated independently, and the correlation between the x and y components is ignored or assumed to be negligible. But, in the vector field, the correlation between the x and y components is retained by treating the vector field as a multivector object.
What follows is a description of Clifford algebras that form the foundation for treating vector fields as multivector objects.
2,0 0,2 3,0 p,q 1 p+q n n We introduce important mathematical concepts and discuss three prominent Clifford algebras, Cl(R), Cl(R), Cl(R), which are directly related to the Clifford neural layers. A more detailed introduction as well as connections to complex numbers and quaternions is provided elsewhere. Consider the vector space Rwith standard Euclidean product <., .>, where n=p+q, and p and q are non-negative integers. A real Clifford algebra Cl(R) is an associative algebra generated by p+q orthonormal basis elements e, . . . , eof the generating vector space R, such that the following quadratic relations hold:
p,q p,q 0,0 0,1 1 1 0,2 1 2 1 2 1 2 1 2 0,2 p+q n 1 2 The pair (p, q) is called the signature and defines a Clifford algebra Cl(R), together with the basis elements that span the vector space Gof Cl(R). Vector spaces of Clifford algebras have scalar elements and vector elements but can also have elements consisting of multiple basis elements of the generating vector space R, which can be interpreted as plane and volume segments. Exemplary low-dimensional Clifford algebras are: (i) Cl(R) which is a one-dimensional algebra that is spanned by the basis element {1} and is therefore isomorphic to R, the field of real numbers; (ii) Cl(R) which is a two-dimensional algebra with vector space Gspanned by {1, e} where the basis vector esquares to 1, and is therefore isomorphic to C, the field of complex numbers; (iii) Cl(R) which is a 4-dimensional algebra with vector space Gspanned by {1, e, e, ee}, where {e, e, ee}all square to 1 and anti-commute. Thus, Cl(R) is isomorphic to the quaternions H.
1 2 1 2 0,2 0 1 2 p,q 0 1 p+q 1 2 1 2 3 2 p+q p+q 2 3 The grade of a Clifford algebra basis element is the dimension of the subspace it represents. For example, the basis elements {1, e, e, ee} of the vector space Gof the Clifford algebra Cl(R) have the grades {0, 1, 1, 2}. Using the concept of grades, one can divide Clifford algebras into linear subspaces made up of elements of each grade. The grade subspace of smallest dimension is M, the subspace of all scalars (elements with 0 basis vectors of the generating vector space). Elements of Mare called vectors, elements of Mare bivectors, and so on. In general, a vector space Gof a Clifford algebra Cl(R) can be written as the direct sum of all of these subspaces: G=M⊕M⊕ . . . ⊕M. The elements of a Clifford algebra are called multivectors, containing elements of subspaces, i.e. scalars, vectors, bivectors, . . . , k-vectors. The basis element with the highest grade is called the pseudoscalar, which in Rcorresponds to the bivector ee, and in Rto the trivector eee.
p+q p+q 2 3 The dual a* of a multivector a is defined as a*=ai, where irepresents the respective pseudoscalar of the Clifford algebra. This definition allows one to relate different multivectors to each other, which is a useful property when defining Clifford Fourier transforms. For example, for Clifford algebras in Rthe dual of the scalar is the bivector, and in R, the dual of the scalar is the trivector.
p+q p+q The geometric product is a bilinear operation on multivectors. For arbitrary multivectors a, b, c∈G, and scalar λ, the geometric product has the following properties: (i) closure, i.e. ab G, (ii) associativity, i.e. (ab)c=a(bc); (iii) commutative scalar multiplication, i.e. λR=aλ; (iv) distributive over addition, i.e. a(b+c)=ab+ac. The geometric product is in general non-commutative, i.e. ab≠ba. Note that Equations 1, 2, and 3 describe the geometric product specifically between basis elements of the generating vector space.
2,0 0,2 1 2 1 2 1 2 1 2 2,0 0,2 2,0 0 1 1 2 2 12 1 2 0 1 1 2 2 12 1 2 Clifford algebras Cl(R) and Cl(R) are now discussed. The 4-dimensional vector spaces of these Clifford algebras have the basis vectors {1, e, e, ee} where e, e, eesquare to +1 for Cl(R) and to −1 for Cl(R). For Cl(R), the geometric product of two multivectors a=a+ae+ae+aeeand b=b+be+be+beeis given by:
1 2 1 1 2 2 2 2 2 2 2 which can be derived by collecting terms that multiply the same basis elements. A vector x=(x, x)∈Rwith standard Euclidean product <., .< can be related to xe+xe∈R⊂G. Clifford multiplication of two vectors x, y∈R⊂Gyields the geometric product xy:
where ∧ is the exterior or wedge product. The asymmetric quantity x∧y=y∧x is associated with the bivector, which can be interpreted as an oriented plane segment.
2 1 2 A unit bivector i, spanned by the (orthonormal) basis vectors eand eis determined by the product:
which if squared yields
2 2 1 2 2 1 1 2 2 2 2 Thus, irepresents a geometric √{square root over (−1)}. From Equation 6, it follows that e=ei=−ieand e=ie−ei.
2 1 2 1 2 Using the pseudoscalar i, the dual of a scalar is a bivector and the dual of a vector is again a vector. The dual pairs of the base vectors are 1↔eeand e↔e. These dual pairs allow one to represent an arbitrary multivector a as
2 2 1 1 2 1 1 2 1 2 1 2 1 which can be regarded as two complex-valued parts: the spinor part, which commutes with the base element 1, i.e. 1i=i1, and the vector part, which anti-commutes with the respective base element e, i.e. ei=eee=−eee=−ie. This decomposition will be the basis for Clifford Fourier transform layers which naturally act on dual pairs.
0,2 0,2 1 2 1 2 The Clifford algebra Cl(R) is isomorphic to the quaternions H, which are an extension of complex numbers and are commonly written in the literature as a+bî+cĵ+d{circumflex over (k)}. Quaternions also form a 4-dimensional algebra spanned by {1, î, ĵ, {circumflex over (k)}}, where î, ĵ, {circumflex over (k)} all square to −1. The algebra isomorphism to Cl(R) is easy to verify since e, e, eeall square to −1 and anti-commute. The basis element 1 is often called the scalar part, and the basis elements î, ĵ, {circumflex over (k)} are called the vector part of a quaternion. Quaternions have practical uses in applied mathematics, particularly for calculations involving three-dimensional rotations, which is used to define the rotational Clifford convolution layer elsewhere.
3,0 3,0 1 2 3 1 2 3 1 2 3 1 2 3 1 2 3 1 2 3 1 2 3 1 2 3 3 2,0 3,0 1 2 3 3 1 1 2 2 3 1 3 1 2 3 Clifford algebra Cl(R) is now discussed. The 8-dimensional vector space Gof the Clifford algebra Cl(R) has the basis vectors {1, e, e, e, ee, ee, ee, eee}, i.e. it consists of one scalar, three vectors {e, e, e}, three bivectors {ee, ee, ee}, and one trivector eee. The trivector is the pseudoscalar iof the algebra. The geometric product of two multivectors is defined analogously to the geometric product of Cl(R). The dual pairs of Cl(R) are: 1↔eee=i, e↔ee, e↔ee, and e↔ee.
3,0 3 x 1 y 2 z 3 x 1 y 2 z 3 3 x 1 x 1 3 x 1 1 2 3 x 2 3 x 1 3 An intriguing example of the duality of the multivectors of Cl(R) emerges when writing the expression of the electromagnetic field F in terms of an electric vector field E and a magnetic vector field B, such that F=E+B, where E=Ee+Ee+Eeand B=Be+Be+Be. In this way the electromagnetic field F decomposes into electric vector and magnetic bivector parts via the pseudoscalar i. For example, for the base component Beof B it holds that Bei=Beeee=Beewhich is a bivector and the dual to the base component Eeof E. Consequently, the multivector representing F consists of three vectors (the electric field components) and three bivectors (the magnetic field components multiplied by i). This viewpoint gives Clifford neural layers a natural advantage over their corresponding scalar DL counterparts.
Clifford convolution layers and Clifford Fourier transform layers for 2-dimensional problems are now described. Extensions to 3 dimensions are provided elsewhere, along with discussions about appropriate normalization and initialization procedures.
2 cin out From regular to Clifford CNN layers. Regular convolutional neural network (CNN) layers take as input feature maps ƒ: Z→Rand convolve them with a set of cfilters
2 2 j i,j j i,j 2 c 2 2 c out out out Equation 8 can be interpreted as an inner product of input feature maps with the corresponding filters at every point y∈Z. By applying cfilters, the output feature maps can be interpreted as cdimensional feature vectors at every point y∈Z. It is desired to extend convolution layers such that the element-wise product of scalars ƒ(y)w(y−x) is replaced by the geometric product of multivector inputs and multivector filters ƒ(y)w(y−x). Therefore, the feature maps ƒ: Z→Rin are replaced by multivector feature maps ƒ: Z→(G)in and convolved with a set of cmultivector filters
out in Note that each (geometric) product, indexed by i∈{1, . . . , c} and j∈{1, . . . c}, now results in a new multivector rather than a scalar. Hence, the output of a layer is a grid of coat multivectors. One can implement a 2D Clifford CNN layer using Equation 4 where
correspond to 4 different kernels representing one 2D multivector kernel, i.e. 4 different convolution layers, and
6 FIG. correspond to the scalar, vector and bivector parts of the input multivector field. The channels of the different layers represent different stacks of scalars, vectors, and bivectors. Analogously, one can implement a 3D Clifford CNN layer using Equation 45. A schematic sketch of a Clifford convolution layer is shown in.
6 FIG. 6 FIG. 6 FIG. 600 600 660 662 664 666 664 illustrates, by way of example, a diagram of an embodiment of a Clifford neural layer. The Clifford neural layerconvolves a multivector inputwith multivector kernels. The convolution can be performed using the geometric product. The result is an output that includes multivector output fields. The geometric productis described elsewhere. For a standard convolution implementation consistent with a convolution of, there would be 8 individual channels, and 32 samples for each channel. Inin contrast, there are only two channels, each with four entry inputs (e.g., scalar, vector, bi-vector, etc. parts of the multivector). A geometric product multiplies the multivector kernel to the entries maintaining dependencies on entries of the multivector inputs.
What follows is pseudocode for a 2D Clifford kernel and a Clifford Convolution layer:
Function CLIFFORDKERNEL2D(W) return KERNEL Function CLIFFORDCONV2D(W, x) KERNEL ← CLIFFORDKERNEL2D(W) Input ← VIEW_AS_REALVECTOR(x) Output ← CLIFFORDCONV(KERNEL, Input) return VIEW_AS_MULTIVECTOR (Output)
0,2 j i,j An alternative parameterization to the Clifford CNN layer introduced in Equation 9, is called a rotational Clifford CNN layer, and can be realized by using the isomorphism of the Clifford algebra Cl(R) to quaternions. Note that a quaternion rotation can be realized by a matrix multiplication. Using the isomorphism, one can represent the feature maps ƒand filters was quaternions:
j i,j j i,j j i,j −1 j i,j j i,j i,j j i,j i,j One can now devise an alternative parameterization of the product between the feature map ƒand w. To be more precise, consider a composite operation that results in a scalar quantity and a quaternion rotation, where the latter acts on the vector part of the quaternion ƒand only produces nonzero expansion coefficients for the vector part of the quaternion output. A quaternion rotation wƒ(w)acts on the vector part (î, ĵ, {circumflex over (k)}) of ƒand can be algebraically manipulated into a vector matrix operation Rƒ, where R: H→H is built up from the elements of w. In other words, one can transform the vector part (î, ĵ, {circumflex over (k)}) of ƒ∈H via a rotation matrix Rthat is built from the scalar and vector part (1, î, ĵ, {circumflex over (k)}) of w∈H. Altogether, a rotational multivector filter
j acts on the feature map ƒthrough a rotational transformation
2 2 c (G)in, and an additional scalar response of the multivector filters: acting on vector and bivector parts of the multivector feature map ƒ: Z
is the scalar output of Equation 4.
i,j A detailed description of the rotational multivector filters R(y x) is outlined elsewhere. While in principle the Clifford CNN layer in Equation 9 and the rotational Clifford CNN layer in Equation 10 are equally flexible, the experiments described below show that rotational Clifford CNN layers lead to better performance.
What follows is pseudocode for a 2D Rotational Clifford kernel and a Rotational Clifford Convolution layer:
layer:
ROT Function CLIFFORDKERNEL 2 D(W) 12 2 2 sq< W [1]+ W [2] 13 2 2 sq< W [1]+ W [3] 23 2 2 sq< W [2]+ W [3] 2 2 2 2 sumsq < W [0]+ W [1]+ W [2]+ W [3]+ ϵ 12 rot← W [0]W [1]/sumsq 13 rot← W [0]W [2]/sumsq 14 rot← W [0]W [3]/sumsq 23 rot← W [1]W [2]/sumsq 24 rot← W [1]W [3]/sumsq 34 rot← W [2]W [3]/sumsq return KERNEL ROT Function CLIFFORDCONV2D(W, x) ROT KERNEL ← CLIFFORDKERNEL2D(W) Input ← VIEW_AS_REALVECTOR(x) Output ← CLIFFORDCONV2D(KERNEL, Input) return VIEW_AS_MULTIVECTOR (Output)
Clifford convolutions satisfy the property of equivariance under translation of the multivector inputs. Analogous to a Theorem presented below, translation equivariance can also be shown for rotational Clifford CNN layers. However, the current definition of Clifford convolutions is not equivariant under multivector rotations or reflections. A general kernel constraint for multivector kernels, which sketches a path towards generalized Clifford convolutions which are equivariant w.r.t rotations or reflections is provided below.
1 n 1 n n From Fourier Neural Operator (FNO) to Clifford Fourier layers. In arbitrary dimension n, the discrete Fourier transform of an n-dimensional complex signal ƒ(x)=ƒ(x, . . . , x): R→C at M× . . . ×Mgrid points is defined as:
in out In Fourier Neural Operators (FNO), discrete Fourier transforms on real-valued input fields and respective back-transforms— implemented as Fast Fourier Transforms on real-valued inputs (RFFTs) are interleaved with a weight multiplication by a complex weight matrix of shape c×cfor each mode, which results in a complex-valued weight tensor of the form
where Fourier modes above cut-off frequencies
are set to zero. Additionally, a residual connection is applied, which is usually implemented as a convolution layer with kernel size 1.
7 FIG. 700 700 770 772 774 778 776 780 w. illustrates, by way of example, a diagram of an embodiment of a Clifford Fourier Neural Operator (CFNO). In the CFNO, a real valued Fast Fourier transform (RFFT) over real valued scalar input fields ƒ (x) is replaced by a complex Fast Fourier transform (FFT) over complex valued dual parts v(x)and s(x)of multivector fields ƒ (x). Pointwise multiplication in the Fourier space via complex weight tensor W is replaced by a geometric productin the Clifford Fourier space via multivector weight tensor W. Additionally, the convolution path is replaced by Clifford convolutions with multivector kernels
7 FIG. 7 FIG. For a standard Fourier transform implementation consistent with a convolution of, there would be 8 individual channels, and 32×32 samples for each channel. Thus, 8 individual Fourier transforms would be performed and the results would be combined. Inin contrast, there are only two channels, each with four entry inputs (e.g., scalar, vector, bi-vector, etc. parts of the multivector). Four Fourier transforms are performed, two Fourier transforms per multivector input.
2 2 2 The 2-dimensional Clifford Fourier transform and the respective inverse transform for multivector valued functions ƒ (x): R→Gand vectors x, ξ∈Rare defined as:
774 2 1 2 provided that the integrals exist. In contrast to standard Fourier transforms, ƒ(x)and {circumflex over (ƒ)}(ξ) represent multivector fields in the spatial and the frequency domain, respectively. Furthermore, i=eeis used in the exponent. Inserting the definition of multivector fields, Equation 12 becomes:
0 0 12 2 1 1 2 2 0 1 2 2 774 776 A Clifford Fourier transform is obtained by applying two standard Fourier transforms to the dual pairs ƒ=ƒ(x)+ƒ(x)iand ƒ=ƒ(x)+ƒ(x)i, which both can be treated as a complex-valued signals ƒ,ƒ:R→C. Consequently, ƒ(x)can be understood as an element of C. The 2D Clifford Fourier transform is the linear combination of two classical Fourier transforms. Discrete versions of Equation 14 are obtained analogously to Equation 11. Similar to FNO, multivector weight tensors
are applied, where again Fourier modes above cut-off frequencies
are set to zero. Thus, the Clifford Fourier modes are point-wise modified via the geometric product as follows:
780 600 k 2,0 7 FIG. The Clifford Fourier modes follow when combining spinor and vector parts of Equation 14. Finally, the residual connection of a prior FNO is replaced by a Clifford convolution with multivector kernel, such as by a Cl(R) convolution layer. A schematic sketch of a Clifford Fourier layer is shown in. 3D Clifford Fourier transforms follow a similar elegant construction, where four separate Fourier transforms to the complex fields as in Equation 16:
770 772 782 784 786 788 786 790 700 In Equation 16 scalar/trivector and vector/bivector components are combined into complex fields and then subjected to a Fourier transform. More details on 2D and 3D Fourier transforms and their properties can be found elsewhere. Spatial representations of v(x)and s(x)are recovered by an inverse FFT which provides v*(x)and s*(x)that are combined to form ƒ*(x). A resultof the Clifford convolution is combined with ƒ*(x)to from an outputof the Clifford FNO.
What follows is pseudocode for implementing a Clifford Spectral Convolution and a Clifford Fourier Layer:
1 2 Function CLIFFORDSPECTRALCONV2D(W, x, m, m) v s x, x← VIEW_AS_DUAL_PARTS(x) v v f(x) ← FFT2(x) - Complex 2D FFT of vector part s s f(x) ← FFT2(x) - Complex 2D FFT of scalar part s v v s f*(x) ← f*(x) · r + f*(x) + f*(x) · i + f*(x) · i − Multivector Fourier modes f*(x) ← f*(x)W − Geometric product in Fourier space v {circumflex over (x)}← IFFT2({circumflex over (f)}*(x)[1] + {circumflex over (f)}*(x)[2]) − Inverse 2D FFT of vector part s {circumflex over (x)}← IFFT2({circumflex over (f)}*(x)[0] + {circumflex over (f)}*(x)[3]) − Inverse 2D FFT of scalar part v s {circumflex over (x)} ← VIEW_AS_MULTIVECTOR({circumflex over (x)}, {circumflex over (x)}) return {circumflex over (x)} f c Function CLIFFORDFOURIERLAYER(W, W, x) 1 f 1 2 y← CLIFFORDSPECTRALCONV(W, x, m, m) 2 x← VIEW_AS_REALVECTOR(x) 2 c 2 y← CLIFFORDCONV(W, x) 2 2 y← VIEW_AS_MULTIVECTOR(y) 1 2 out ← ACTIVATION (y+ y) return out
5 8 FIG. An assessment of Clifford neural layers was performed for different architectures in three experimental settings: the incompressible Navier-Stokes equations, shallow water equations for weather modeling, and 3-dimensional Maxwell's equations in matter. Ablations were also performed for different dataset sizes. All datasets share the common trait of containing multiple input and output fields. More precisely, one scalar and one 2-dimensional vector field in case of the Navier-Stokes and the shallow water equations, and a 3-dimensional (electric) vector field and its dual (magnetic) bivector field in case of the Maxwell's equations. ResNet and Fourier Neural Operator (FNO) based architectures were replaced with their respective Clifford algebra neural layer counterparts. In doing so, every convolution and every Fourier transform was changed to a respective Clifford neural layer operation. Also, a substitution for normalization techniques and activation functions with appropriate counterparts was performed. Different training set sizes were evaluated, and losses for scalar and vector fields are reported. Inputs to the neural PDE surrogates are respective fields at previous t timesteps, where t varies between different PDEs. The one-step loss is the mean-squared error at the next timestep summed over fields. The rollout loss is the mean-squared error after applying the neural PDE surrogatetimes, summing over fields and time dimension. More information on the datasets, implementation details of the tested architectures, loss functions, and more detailed results can be found elsewhere.shows example rollouts of the scalar and the vector field components obtained by Clifford Fourier PDE surrogates for a qualitative comparison.
The summed MSE (SMSE) loss used in these experiments are defined as:
fields y SMSE t t fields One-step loss where N=1 and Ncomprises all scalar and vector components. t fields Vector loss where N=1 and Ncomp rises only vector components. t fields Scalar loss where N=1 and Ncomprises only the scalar field. t fields Rollout loss where N=5 and Ncomprises all scalar and vector components. where u is the target, û the model output, Ncomprises scalar fields as well as individual vector field components, and Nis the total number of spatial points. The LEquation is used for training with N=1, and four metrics:
For Maxwell's equation, electric and magnetic loss are defined analogously to the vector and the scalar loss for Navier-Stokes and shallow water experiments.
Experiments are performed with two architecture families: ResNet models and Fourier Neural Operators (FNOs). All baseline models are fine-tuned for all individual experiments with respect to number of blocks, number of channels, number of modes (FNO), learning rates, normalization and initialization procedures, and activation functions. The best models are reported, and for reported Clifford results each convolution layer is substituted with a Clifford convolution, each Fourier layer with a Clifford Fourier layer, each normalization with a Clifford normalization and each non-linearity with a Clifford non-linearity. A Clifford non-linearity in this context is the application of the corresponding default linearity to the different multivector components.
For Navier-Stokes and shallow water experiments, ResNet architectures with 8 residual blocks, each consisting of two convolution layers with 3×3 kernels, shortcut connections, group normalization, and Gaussian Error Linear Unit (GeLU) activation functions were used. Two embedding and two output layers were used, i.e. the overall architectures could be classified as Res-20 networks. In contrast to standard residual networks for image classification, down-projection techniques were not used. Down projections are convolution layers with strides larger than 1 or via pooling layers. In contrast, the spatial resolution stays constant throughout the network. The same number of hidden channels throughout the network, that is 128 channels per layer were used. Overall this results in roughly 1.6 million parameters. Increasing the number of residual blocks or the number of channels did not increase the performance significantly.
For every ResNet-based experiment, the fine-tuned ResNet architectures were replaced with two Clifford counterparts: each CNN layer is replaced with a (i) Clifford CNN layer, and (ii) with a rotational Clifford CNN layer. To keep the number of weights similar, instead of 128 channels the resulting architectures have 64 multivector channels, resulting again in roughly 1.6 million floating point parameters. Additionally for both architectures, GeLU activation functions are replaced with Clifford GeLU activation functions, group normalization is replaced with Clifford group normalization. Using Clifford initialization techniques did not improve results.
For Navier-Stokes and shallow water experiments, 2-dimensional Fourier Neural Operators (FNOs) consisting of 8 FNO blocks, two embedding and two output layers were used. Each FNO block comprised a convolution path with a 1×1 kernel and an FFT path. 16 Fourier modes (for x and y components) for point-wise weight multiplication were used, and overall use 128 hidden channels. GeLU activation functions were used. Additional shortcut connections or normalization techniques, such as batchnorm or group were used. Note norm did not improve performance, neither did larger numbers of hidden channels, nor more FNO blocks. Overall this resulted in roughly 140 million parameters for FNO based architectures.
For 3-dimensional Maxwell experiments, 3-dimensional Fourier Neural Operators (FNOs) consisting of 4 FNO blocks, two embedding and two output layers were used. Each FNO block comprised a 3D convolution path with a 1×1 kernel and an FFT path. 6 Fourier modes (for x, y, and z components) for point-wise weight multiplication were used, and overall used 96 hidden channels. Interestingly, using more layers or more Fourier modes degraded performances. Similar to the 2D experiments, GeLU activation functions were used, and did not apply shortcut connections nor normalization techniques, such as batchnorm or groupnorms. Overall this resulted in roughly 65 million floating point parameters for FNO based architectures.
For every FNO-based experiment, the fine-tuned FNO architectures were replaces with respective Clifford counterparts: each FNO layer was replaced by its Clifford counterpart. To keep the number of weights similar, instead of 128 channels the resulting architectures have 48 multivector channels, resulting in roughly the same number of parameters. Additionally, GeLU activation functions are replaced with Clifford GeLU activation functions. Using Clifford initialization techniques did not improve results.
For 3-dimensional Maxwell experiments, each 3D Fourier transform layer was replaced with a 3D Clifford Fourier layer and each 3D convolution with a respective Clifford convolution. 6 Fourier modes (for x, y, and z components) for point-wise weight multiplication were used, and overall used 32 hidden multivector channels, which results in roughly the same number of parameters (55 million). In contrast to 2-dimensional implementations, Clifford initialization techniques proved important for 3-dimensional architectures. Most notably, too large initial values of the weights of Clifford convolution layers hindered gradient flows through the Clifford Fourier operations.
−4 −4 −4 SMSE Models were optimized using the Adam optimizer with learning rates [10, 2·10, 5 10] for 50 epochs and minimized the summed mean squared error (SMSE) which is outlined in LEquation. Cosine annealing was used as a learning rate scheduler with a linear warmup. For baseline ResNet models, the number of layers, number of channels, and normalization procedures were optimized. Different activation functions were tested. For baseline FNO models, number of layers, number of channels, and number of Fourier modes were optimized. Larger numbers of layers or channels did not improve the performances for both ResNet and FNO models. For the respective Clifford counterparts, convolution and Fourier layers were replaced by Clifford convolution and Clifford Fourier layers, respectively. Clifford normalization schemes were used. The number of layers was adjusted to obtain similar numbers of parameters. Clifford architectures could have been optimized slightly more by, for example, using different numbers of hidden layers than the baseline models did. However, this would (i) slightly be against the argument of having “plug- and play” replace layers, and (ii) would have added quite some computational overhead.
All FNO and CFNO experiments used 4×16 GB NVIDIA V100 machines for training. All ResNet and Clifford ResNet experiments used 8×32 GB NVIDIA V100 machines. Average training times varied between 3 h and 48 h, depending on task and number of trajectories. Clifford runs on average took twice as long to train for equivalent architectures and epochs.
The incompressible Navier-Stokes equations are built upon momentum and mass conservation of fluids expressed for the velocity flow field v:
2 9 FIG. where v·∇v is the convection, μ∇v the viscosity, ∇p the internal pressure and ƒ an external force. In addition to the velocity field v(x), a scalar field s(x) can be introduced that represents a scalar quantity, for example, smoke, that is being transported via the velocity field. The scalar field is advected by the vector field, as the vector field changes, the scalar field is transported along with it, whereas the scalar field influences the vector field only via the external force term ƒ. This is called a weak coupling between vector and scalar fields. The 2D Navier-Stokes equation can be modeled using ΦFlow, obtaining data on a grid with spatial resolution of 128×128 (Δx=0.25, Δy=0.25), and temporal resolution of Δt=1.5 s. Results for one-step loss and rollout loss on the test set are shown in.
Convection is the rate of change of a vector field along a vector field (in this case along itself), viscosity is the diffusion of a vector field, i.e. the net movement form higher valued regions to lower concentration regions, μ is the viscosity coefficient. The incompressibility constrained yields mass conservation via ∇·v=0. For example, v might represent velocity of air inside a room, and s might represent concentration of smoke. Similar to convection, advection is the transport of a scalar field along a vector field
The 2D Navier-Stokes equation can be implemented using ΦFlow. Solutions are propagated in which one can solve for the pressure field and subtract its spatial gradients afterwards. Semi-Lagrangian advection (convection) is used for v, and McCormack advection for s. Additionally, express the external buoyancy force ƒ in Equation 17 as force acting on the scalar field. Solutions are obtained using Boussinesq approximation, which ignores density differences except where they appear in terms multiplied by the acceleration due to gravity. The essence of the Boussinesq approximation is that the difference in inertia is negligible but gravity is sufficiently strong to make the specific weight appreciably different between the two fluids.
9 FIG. rot rot illustrates, by way of example, graphs of one-step loss and rollout loss for a variety of NN PDE modeling architectures implementing Navier-Stokes equation. For ResNet-like architectures, observe that both CResNet (Clifford ResNet or ResNet with layers replaced with Clifford neural layers) and CResNetrot (rotational Clifford ResNet or ResNet with layers replaced with rotational Clifford neural layers) improve upon the ResNet baseline. Additionally, observe that rollout losses are also lower for the two Clifford based architectures, which can be attributed to better and more stable models that do not overfit to one-step predictions so easily. Lastly, while in principle CResNet and CResNetbased architectures are equally flexible, CResNetin general performs better than CResNet in this scenario. For FNO and respective Clifford Fourier based (CFNO) architectures, the loss is in general much lower than for ResNet based architectures. CFNO architectures improve upon FNO architectures for all dataset sizes, and for one-step as well as rollout losses.
TABLE 1 Model comparison on four different metrics for neural PDE surrogates which are trained on Navier-Stokes training datasets of varying size. Error bars are obtained by running experiments with three different initial seeds. SMSE Method Trajs. scalar vector onestep rollout ResNet 0.00300 ± 0.00003 0.01255 ± 0.00008 0.01553 ± 0.00011 0.11362 ± 0.00048 CResNet 2080 0.00566 ± 0.00062 0.02252 ± 0.00284 0.02806 ± 0.00346 0.15844 ± 0.02677 rot CResNet 0.00376 ± 0.00028 0.01413 ± 0.00116 0.01780 ± 0.00143 0.10681 ± 0.00476 ResNet 0.00341 ± 0.00028 0.01431 ± 0.00102 0.01767 ± 0.00138 0.13234 ± 0.00020 CResNet 5200 0.00265 ± 0.00004 0.00988 ± 0.00024 0.01250 ± 0.00022 0.09975 ± 0.00060 rot CResNet 0.00234 ± 0.00014 0.00857 ± 0.00066 0.01087 ± 0.00074 0.09427 ± 0.00071 ResNet 0.00321 ± 0.00004 0.01337 ± 0.00044 0.01653 ± 0.00048 0.13802 ± 0.00223 CResNet 10400 0.00315 ± 0.00006 0.01162 ± 0.00019 0.01473 ± 0.00018 0.10671 ± 0.00246 rot CResNet 0.00201 ± 0.00020 0.00719 ± 0.00074 0.00917 ± 0.00090 0.10005 ± 0.00229 ResNet 0.00342 ± 0.00003 0.01379 ± 0.00079 0.01716 ± 0.00091 0.13030 ± 0.00379 CResNet 15600 0.00285 ± 0.00019 0.01076 ± 0.00051 0.01357 ± 0.00063 0.10372 ± 0.00269 rot CResNet 0.00204 ± 0.00014 0.00736 ± 0.00069 0.00938 ± 0.00087 0.09799 ± 0.00139 FNO 2080 0.00318 ± 0.00021 0.00613 ± 0.00044 0.00931 ± 0.00064 0.04281 ± 0.00300 CFNO 0.00266 ± 0.00002 0.00484 ± 0.00006 0.00749 ± 0.00008 0.03461 ± 0.00031 FNO 5200 0.00204 ± 0.00004 0.00332 ± 0.00011 0.00536 ± 0.00015 0.02684 ± 0.00067 CFNO 0.00189 ± 0.00001 0.00293 ± 0.00002 0.00482 ± 0.00003 0.02430 ± 0.00012 FNO 10400 0.00156 ± 0.00003 0.00220 ± 0.00007 0.00375 ± 0.00010 0.02005 ± 0.00042 CFNO 0.00148 ± 0.00001 0.00205 ± 0.00001 0.00353 ± 0.00002 0.01886 ± 0.00006
x y The shallow water equations describe a thin layer of fluid of constant density in hydrostatic balance, bounded from below by the bottom topography and from above by a free surface. For example, the deep water propagation of a tsunami can be described by the shallow water equations. For simplified weather modeling, the shallow water equations express the velocity in x-direction, or zonal velocity v, the velocity in the y-direction, or meridional velocity v, the acceleration due to gravity g, and the vertical displacement of free surface η(x,y), which subsequently is used to derive pressure fields:
10 FIG. The topography of the earth is modeled by h(x, y). The relation between vector and scalar components is relatively strong (strong coupling between vector and scalar fields due to the 3-coupled PDEs of Equation 18). One can modify the implementation in SpeedyWeather.jl, obtaining data on a grid with spatial resolution of 192×96 (Δx=1.875°, Δy=3.75°), and temporal resolution of Δt=6 h. Results for one-step loss and rollout loss on the test set are shown in. The equation is solved on a closed domain with periodic boundary conditions. The simulation was rolled out for 20 days and sample every 6 h. Here 20 days is of course not the actual simulation time but rather the simulated time. Trajectories contain scalar pressure and wind vector fields at 84 different time points.
SpeedyWeather.jl combines the shallow water equations with spherical harmonics for the linear terms and Gaussian grid for the non-linear terms with the appropriate spectral transforms. It internally uses a leapfrog time scheme with a Robert and William's filter to dampen the computational modes and achieve 3rd order accuracy. SpeedyWeather.jl is based on the atmospheric general circulation model SPEEDY in Fortran.
10 FIG. rot rot illustrates, by way of example, MSE loss for one-step and rollout for various NN architectures implementing Navier-Stokes equations. Observe similar performance properties as for the Navier-Stokes experiments discussed previously. However, performance differences between baseline and Clifford architectures are even more pronounced, which can be attributed to the stronger coupling of the scalar and the vector fields. For ResNet-like architectures, CResNet and CResNetimprove upon the ResNet baseline, rollout losses are much lower for the two Clifford based architectures, and CResNetbased architectures in general perform better than CResNet based ones. For Fourier based architectures, the loss is in general much lower than for ResNet based architectures (a training set size of 56 trajectories yields similar (C)FNO test set performance than a training set size of 896 trajectories for ResNet based architectures). CFNO architectures improve upon FNO architectures for all dataset sizes, and for one-step as well as rollout losses, which is especially pronounced for low number of training trajectories.
TABLE 2 Model comparison on four different metrics for neural PDE surrogates which are trained on the shallow water equations training datasets of varying size. Results are obtained by using a two timestep history input. Error bars are obtained by running experiments with three different initial seeds. SMSE Method Trajs. scalar vector onestep rollout ResNet 0.0240 ± 0.0002 0.0421 ± 0.0010 0.0661 ± 0.0011 1.1195 ± 0.0197 CResNet 192 0.0617 ± 0.0016 0.0823 ± 0.0027 0.1440 ± 0.0042 2.0423 ± 0.0494 rot CResNet 0.0319 ± 0.0003 0.0576 ± 0.0005 0.0894 ± 0.0007 1.4756 ± 0.0044 ResNet 0.0140 ± 0.0003 0.0245 ± 0.0007 0.0385 ± 0.0010 0.7083 ± 0.0119 CResNet 448 0.0238 ± 0.0007 0.0448 ± 0.0023 0.0685 ± 0.0030 1.1727 ± 0.0483 rot CResNet 0.0114 ± 0.0001 0.0221 ± 0.0001 0.0335 ± 0.0002 0.6127 ± 0.0018 ResNet 0.0086 ± 0.0000 0.0156 ± 0.0003 0.0242 ± 0.0003 0.4904 ± 0.0080 CResNet 896 0.0095 ± 0.0000 0.0183 ± 0.0004 0.0278 ± 0.0006 0.5247 ± 0.0101 rot CResNet 0.0055 ± 0.0000 0.0106 ± 0.0001 0.0161 ± 0.0001 0.3096 ± 0.0010 ResNet 0.0061 ± 0.0002 0.0123 ± 0.0009 0.0184 ± 0.0010 0.4780 ± 0.0062 CResNet 1792 0.0039 ± 0.0000 0.0071 ± 0.0000 0.0111 ± 0.0001 0.2842 ± 0.0067 rot CResNet 0.0025 ± 0.0000 0.0044 ± 0.0000 0.0069 ± 0.0000 0.2370 ± 0.0000 ResNet 0.0060 ± 0.0002 0.0121 ± 0.0003 0.0181 ± 0.0005 0.4480 ± 0.0058 CResNet 2048 0.0039 ± 0.0001 0.0072 ± 0.0002 0.0111 ± 0.0003 0.2816 ± 0.0065 rot CResNet 0.0028 ± 0.0005 0.0480 ± 0.0006 0.0075 ± 0.0011 0.2164 ± 0.0070 FNO 56 0.0271 ± 0.0016 0.0345 ± 0.0007 0.0616 ± 0.0022 0.8032 ± 0.0043 CFNO 0.0071 ± 0.0003 0.0177 ± 0.0004 0.0250 ± 0.0007 0.4323 ± 0.0046 FNO 192 0.0021 ± 0.0002 0.0057 ± 0.0001 0.0077 ± 0.0003 0.1444 ± 0.0026 CFNO 0.0012 ± 0.0000 0.0040 ± 0.0001 0.0053 ± 0.0001 0.0941 ± 0.0021 FNO 448 0.0007 ± 0.0001 0.0026 ± 0.0000 0.0034 ± 0.0001 0.0651 ± 0.0014 CFNO 0.0005 ± 0.0000 0.0020 ± 0.0000 0.0026 ± 0.0001 0.0455 ± 0.0009 FNO 896 0.0004 ± 0.0000 0.0016 ± 0.0000 0.0020 ± 0.0001 0.0404 ± 0.0005 CFNO 0.0003 ± 0.0000 0.0013 ± 0.0000 0.0017 ± 0.0001 0.0315 ± 0.0004
TABLE 3 Model comparison on four different metrics for neural PDE surrogates which are trained on the shallow water equations training datasets of varying size. Results are obtained by using a four timestep history input. Error bars are obtained by running experiments with three different initial seeds. SMSE Method Trajs. scalar vector onestep rollout FNO 56 0.0276 ± 0.0017 0.0388 ± 0.0023 0.0663 ± 0.0038 0.6821 ± 0.0379 CFNO 0.0093 ± 0.0003 0.0252 ± 0.0005 0.0345 ± 0.0009 0.4357 ± 0.0056 FNO 192 0.0033 ± 0.0007 0.0069 ± 0.0009 0.0102 ± 0.0015 0.1612 ± 0.0057 CFNO 0.0015 ± 0.0001 0.0050 ± 0.0003 0.0065 ± 0.0003 0.1023 ± 0.0026 FNO 448 0.0009 ± 0.0001 0.0023 ± 0.0002 0.0032 ± 0.0003 0.0687 ± 0.0023 CFNO 0.0010 ± 0.0006 0.0039 ± 0.0027 0.0050 ± 0.0033 0.1156 ± 0.0913 FNO 896 0.0004 ± 0.0001 0.0012 ± 0.0001 0.0016 ± 0.0001 0.0436 ± 0.0011 CFNO 0.0003 ± 0.0000 0.0012 ± 0.0001 0.0015 ± 0.0001 0.0353 ± 0.0010
Electromagnetic simulations play a critical role in understanding light-matter interaction and designing optical elements. Neural networks have been already successful applied in inverse-designing photonic structures.
Maxwell's equations in matter are:
0 R 0 r 0 t 0 r In isotropic media, Maxwell's equations in matter propagate solutions of the displacement field D, which is related to the electrical field via D=ϵϵE, where ϵis the permittivity of free space and ϵis the permittivity of the medium, and the magnetization field H, which is related to the magnetic field B via H=μμB, where μis the permeability of free space and μis the permeability of the medium. Lastly, j is the electric current density and ρ the total electric charge density.
3 The electromagnetic field F has the intriguing property that the electric field E and the magnetic field B are dual pairs, thus F=E+Bi. There is a strong coupling between the electric field and its dual (bivector) magnetic field. This duality also holds for D and H. The solution of Maxwell's equation in matter can be propagated using a finite-difference time-domain method, where the discretized Maxwell's equations are solved in a leapfrog manner. First, the electric field vector components in a volume of space are solved at a given instant in time. Second, the magnetic field vector components in the same spatial volume are solved at the next instant in time.
−7 11 FIG. For the experiment, randomly place 18 (6 in the x-y plane, 6 in the x-z plane, 6 in the y-z plane) different light sources outside a cube which emit light with different amplitudes and different phase shifts, causing the resulting D and H fields to interfere with each other. Data was obtained on a grid with spatial resolution of 32×32×32 (Δx=Δy=Δz=5 10m), and temporal resolution of Δt=50 s. FNO-based architectures were tested along with respective Clifford counterparts (CFNO). Note that Maxwell's equations are in general relatively simple due to the wave-like propagation of electric and magnetic fields. It is however a good playground to demonstrate the inductive bias advantages of Clifford based architectures, which leverage the vector-bivector character of electric and magnetic field components. Results for one-step loss and rollout loss on the test set are shown in.
3 3 x 1 x 1 3 x 1 2 3 x 2 3 1 3 3 1 3 2 3 2 3 3 1 3 3 1 2 3 The two fields are strongly coupled, e.g. temporal changes of electric fields induce magnetic fields and vice versa. Probably the most illustrative co-occurrence of electric and magnetic fields is when describing the propagation of light. In standard vector algebra, E is a vector while B is a pseudovector, i.e. the two kinds of fields are distinguished by a difference in sign under space inversion. F=E+Binaturally decomposes the electromagnetic field into vector and bivector parts via the pseudoscalar i. For example, for the base component Beof B it holds that Bei=Beee=Bee, which is a bivector and the dual to the base component eof E. Geometric algebra reveals that a pseudovector is nothing else than a bivector represented by its dual, so the magnetic field B in F=E+Biis fully represented by the complete bivector Bi, rather than B alone. Consequently, the multivector representing F consists of three vectors (the electric field components) and three bivectors ei=ee, ei=ee, ei=ee(the magnetic field components multiplied by i).
11 FIG. illustrates, by way of example, graphs of MSE versus number of training trajectories for FNO and CFNO NN architectures for Maxwell's Equations simulations. Note that CFNO architectures improve upon FNO architectures, especially for low numbers of trajectories. Results demonstrate the much stronger inductive bias of Clifford based 3-dimensional Fourier layers, and their general applicability to 3-dimensional problems, which are structurally even more interesting than 2-dimensional ones.
TABLE 4 Model comparison on four different metrics for neural PDE surrogates which are trained on the Maxwell equations training datasets of varying size. Results are obtained by using a two timestep history input. Error bars are obtained by running experiments with three different initial seeds. SMSE Method Trajs. D H onestep rollout FNO 640 0.0030 ± 0.0006 0.00233 ± 0.00050 0.0054 ± 0.0011 0.0186 ± 0.0083 CFNO 0.0006 ± 0.0001 0.00072 ± 0.00010 0.0013 ± 0.0002 0.0054 ± 0.0023 FNO 1280 0.0010 ± 0.0002 0.00085 ± 0.00020 0.0019 ± 0.0004 0.0068 ± 0.0036 CFNO 0.0003 ± 0.0001 0.00041 ± 0.00010 0.0007 ± 0.0002 0.0029 ± 0.0016 FNO 3200 0.0003 ± 0.0001 0.00025 ± 0.00010 0.0005 ± 0.0001 0.0020 ± 0.0011 CFNO 0.0002 ± 0.0000 0.00020 ± 0.00010 0.0004 ± 0.0001 0.0015 ± 0.0009 FNO 6400 0.0001 ± 0.0000 0.00009 ± 0.00000 0.0002 ± 0.0000 0.0008 ± 0.0004 CFNO 0.0001 ± 0.0000 0.00009 ± 0.00000 0.0002 ± 0.0000 0.0007 ± 0.0004
Embodiments introduce Clifford neural layers. A convolutional Clifford layer, Fourier Clifford layer, rotational Clifford convolutional layer, and an equivariant Clifford layer are provided. The Clifford neural layers handle the various scalar (e.g., charge density), vector (e.g. electric field), bivector (magnetic field) and higher order fields as proper geometric objects organized as multivectors. This geometric algebra perspective, provide for a natural generalizable convolution and Fourier transformation to their Clifford counterparts. Embodiments provide an elegant and natural rule to design new neural network layers. The multivector viewpoint of the Clifford neural layers led to better representation of the relationship between different fields and their individual components, allowing for significant outperformance of equivalent standard architectures for neural PDE surrogates. The performance gap increased with less data availability across all settings, denoting better inductive bias of the Clifford neural layers.
One limitation is the current speed of Fast Fourier Transform (FFT) operations on machine learning accelerators like graphics processing units (GPUs). While an active area of research, current available versions of cuFFT kernels wrapped in PyTorch are not yet as heavily optimized, especially for the gradient pass. Furthermore, FFTs of real-valued signals are Hermitian-symmetric, and outputs contain only positive frequencies below the Nyquist frequency. Such real-valued FFTs are used for Fourier Neural Operator (FNO) layers. Clifford Fourier layers on the other hand use complex-valued FFT operations where the backward pass is approximately twice as slow. For similar parameter counts, inference times of FNO and CFNO networks are similar. The speed of geometric convolution layers was already investigated. On the one hand, Clifford convolutions are more parameter efficient than the standard convolutions since they share parameters among filters, on the other hand, the net number of operations is larger for the same number of parameters resulting in increased training times by a factor of about 2. Finally, from a PDE point of view, the presented approaches to obtain PDE surrogates are limited since the NN have to be retrained for different equation parameters or e.g., different Δt. Custom multivector GPU kernels can overcome many of the speed issues as the compute density of Clifford operations is much higher which is better for hardware accelerators
Besides PDE modeling, weather predictions, electro-magnetic field interactions, and fluid dynamics, applications of multivector representations and Clifford layers can be found in neural implicit representations.
2,0 0,2 3,0 n A closer look at Cl(R), Cl(R), and Cl(R) is now provided. A vector space over a field F is a set V together with two binary operations that satisfy the axioms for vector addition and scalar multiplication. The axioms of addition ensure that if two elements of V get added together, another element of V is provided as a result. The elements of F are called scalars. Examples of a field F are the real numbers R and the complex numbers C. Although it is common practice to refer to the elements of a general vector space V as vectors, to avoid confusion this discussion reserves the usage of this term to the more specific case of elements of R. As shown below, general vector spaces can consist of more complicated, higher-order objects than scalars, vectors or matrices.
1 2 1 2 1 2 1 2 1 2 An algebra over a field consists of a vector space V over a field F together with an additional bilinear law of composition of elements of the vector space, V×V→V, that is, if a and b are any two elements of V, then ab: V×V→V is an element of V, satisfying a pair of distribution laws: a(λb+λc)=λab+λac and (λa+λb)c=λac+λbc for λ, λ∈F and a, b, c∈V. Note that general vector spaces do not have bilinear operations defined on their elements.
n n n n 1 n p,q Focus is now on Clifford algebras over R. A real Clifford algebra is generated by the n-dimensional vector space Rthrough a set of relations that hold for the basis elements of the vector space R. Denote the basis elements of Rwith e, . . . , e, and without loss of generality choose these basis elements to be mutually orthonormal. Take two non-negative integers p and q, such that p+q=n, then a real Clifford algebra Cl(R) with the “signature” (p, q), is generated through the following relations that define how the bilinear product of the algebra operates on the basis elements of R:
i j j i i j k 1 2 p+q Through these relations one can generate a basis for the vector space of the Clifford algebra, which we is denoted by G. Equations 19 and 20 show that the product between two vectors yields a scalar. According to the aforementioned definition of an algebra over a field, a Clifford algebra with a vector space G is equipped with a bilinear product G×G→G, that combines two elements from the vector space G and yields another element of the same space G. Therefore, both scalars and vectors are elements of the vector space G. Equation 21 shows that besides scalar and vector elements, higher order elements including a combination of two basis elements, such as eeand ee, are also part of the vector space G. By combining Equations 19, 20, 21 one can create even higher order elements such as eeefor i≠j≠k, or ee. . . e, which all must be part of the vector space G.
p,q σ(1) σ(2) σ(k) 1 2 k 1 2 p+q 1 2 p+q−1 p+q 1 2 p+q n To determine the basis elements that span the vector space G of Cl(R), note that elements ee. . . eand ee. . . eare related through a simple scalar multiplicative factor of plus or minus one, depending on the sign of the permutation σ. Therefore, it suffices to consider the unordered combinations of basis elements of R: the basis of the vector space G is given by {1, e, e, . . . e, ee, . . . , ee, . . . , ee. . . e}.
n n n n n p+q p,q In summary, two different vector spaces have been introduced. First, the vector space Rwhich generates the Clifford algebra, and second the vector space G, which is the vector space spanned by the basis elements of the Clifford algebra Cl(R). Convention is to denote the vector space of a real Clifford algebra with a superscript n of the dimension of the generating vector space, yielding Gfor a generating vector space R. Note that the dimension of the vector space Gis 2=2.
0,0 0,1 1 1 0,2 1 2 1 2 1 2 1 2 0,2 1 2 Common examples of low-dimensional Clifford algebras are: (i) Cl(R) which is a one-dimensional algebra that is spanned by the vector {1} and is therefore isomorphic to R, the field of real numbers; (ii) Cl(R) which is a two-dimensional algebra with vector space Gspanned by {1, e} where the basis vector esquares to −1, and is therefore isomorphic to C, the field of complex numbers; (iii) Cl(R) which is a 4-dimensional algebra with vector space Gspanned by {1, e, e, ee}, where e, e, eeall square to −1 and anti-commute. Thus, Cl(R) is isomorphic to the quaternions H.
1 2 1 2 0,2 2,0 0 1 2 p,q p+q The grade of a Clifford algebra basis element is the dimension of the subspace it represents. For example, the basis elements {1, e, e, ee} of the Clifford algebras Cl(R) and Cl(R) have the grades {0, 1, 1, 2}. Using the concept of grades, one can divide the vector spaces of Clifford algebras into linear subspaces made up of elements of each grade. The grade subspace of smallest dimension is M, the subspace of all scalars (elements with 0 basis vectors). Elements of Mare called vectors, elements of Mare bivectors, and so on. In general, the vector space Gof a Clifford algebra Clcan be written as the direct sum of all of these subspaces:
2 3 1 2 1 2 3 p+q The elements of a Clifford algebra are called multivectors, containing elements of subspaces (e.g., scalars, vectors, bivectors, trivectors, etc.). The basis element with the highest grade is called the pseudoscalar, which in Rcorresponds to the bivector ee, and in Rto the trivector eee. The pseudoscalar is often denoted with the symbol i.
p + q p + q p + q Using Equations 19, 20, 21, it is seen how basis elements of the vector space Gof the Clifford algebra are formed using basis elements of the generating vector space V. Note how elements of Gare combined, in other words, how multivectors are bilinearly operated on. The geometric product is the bilinear operation on multivectors in Clifford algebras. For arbitrary multivectors a, b, c∈G, and scalar λ the geometric product has the following properties:
The geometric product is in general non-commutative, i.e. ab≠ba. The geometric product is made up of two things: an inner product (that captures similarity) and exterior (wedge) product that captures difference.
The dual a* of a multivector a is defined as:
p+q where irepresents the respective pseudoscalar of the Clifford algebra.
2 3 This definition allows one to relate different multivectors to each other, which is a useful property when defining Clifford Fourier transforms. For example, for Clifford algebras in Rthe dual of a scalar is a bivector, and for the Clifford algebra Rthe dual of a scalar is a trivector.
0,1 1 1 0,1 1 1 0 1 1 0 1 1 1 The Clifford algebra Cl(R) is a two-dimensional algebra with vector space Gspanned by {1, e}, and where the basis vector esquares to −1. Cl(R) is thus algebra-isomorphic to C, the field of complex numbers. This becomes more obvious by identifying that the basis element with the highest grade, i.e. e, as the pseudoscalar iwhich is the imaginary part of the complex numbers. The geometric product between two multivectors a=a+aeand b=b+beis therefore also isomorphic to the product of two complex numbers:
2,0 1 2 1 2 1 2 1 2 0 1 1 2 2 12 1 2 0 1 1 2 2 12 1 2 2 The Clifford algebra Cl(R) is a 4-dimensional algebra with vector space Gspanned by the basis vectors {1, e, e, ee} where e, e, eeall square to +1. The geometric product of two multivectors a=a+ae+ae+aeeand b=b+be+be+beeis defined as:
1 1 2 2 i j j i 1 2 1 2 1 2 Using the relations ee=1, ee=1, and ee=−ee, for i≠j≠{e, e}, from which it follows that eeee=−1, and:
2 2 2 2 1 1 2 2 A vector x∈R⊂Gis identified with xe+xe∈R⊂G. Clifford multiplication of two vectors
2 2 x, y∈R⊂Gyields the geometric product xy:
The asymmetric quantity x∧y=−y∧x is associated with the bivector, which can be interpreted as an oriented plane segment.
Equation 31 can be rewritten to express the (symmetric) inner product and the (anti-symmetric) outer product in terms of the geometric product:
2 2,0 1 2 1 2 1 2 2 1 2 1 2 From the basis vectors of the vector space Gof the Clifford algebra Cl(R), i.e. {1, e, e, ee}, probably the most interesting is ee. A closer look at the unit bivector i=eewhich is the plane spanned by eand eand determined by the geometric product is provided:
1 2 2 where the inner producte, eis zero due to the orthogonality of the base vectors. The bivector i, if squared yields
2 and thus irepresents a true geometric √{square root over (−1)}. From Equation 34, it follows that
2 2 1 2 1 2 The dual of a multivector a∈Gis defined via the bivector as ia. Thus, the dual of a scalar is a bivector and the dual of a vector is again a vector. The dual pairs of the base vectors are 1↔eeand e↔e. These dual pairs allow one to write an arbitrary multivector a as
2 2 which can be regarded as two complex-valued parts: the spinor part, which commutes with iand the vector part, which anti-commutes with i.
0,2 1 2 1 2 1 2 1 2 0,2 2 The Clifford algebra Cl(R) is a 4-dimensional algebra with vector space Gspanned by the basis vectors {1, e, e, ee} where e, e, eeall square to −1. The Clifford algebra Cl(R) is algebra-isomorphic to the quaternions H, which are commonly written as a+bî+cĵ+d{circumflex over (k)}, where the (imaginary) base elements î, ĵ, {circumflex over (k)} fulfill the relations:
0,2 Quaternions also form a 4-dimensional algebra spanned by {1, î, ĵ, {circumflex over (k)}}, where î, ĵ, {circumflex over (k)} all square to −1. The basis element 1 is often called the scalar part, and the basis elements î, ĵ, {circumflex over (k)} are called the vector part of a quaternion. The geometric product of two multivectors in Cl(R) is shown in Equations (29) and (30).
3 1 2 3 1 2 1 3 2 3 1 2 3 1 2 3 1 2 1 3 2 3 1 2 3 3 2,0 The Clifford algebra is an 8-dimensional algebra with vector space Gspanned by the basis vectors {1, e, e, e, ee, ee, ee, eee} i.e. one scalar, three vectors {e, e, e}, three bivectors {ee, ee, ee}, and one trivector eee. The trivector is the pseudoscalar iof the algebra. The geometric product of two multivectors is defined analogously to the geometric product of Cl(R), following the associative and bilinear multiplication of multivectors follows:
3,0 Using Definition 2, the dual pairs of Clare:
3,0 The geometric product for Cl(R) is defined analogously to the geometric product of Cl2,0(R) via:
Where minus signs appear due to reordering of basis elements. Equation 44 simplifies to:
2 2 3 A derivation of the implementation of translation equivariant Clifford convolution layers for multivectors in G, multivectors of Clifford algebras generated by the 2-dimensional vector space R, is now provided. Refer to Equations (8)-(10) for details regarding a Clifford Convolution layer. Then an extension to Clifford algebras generated by the 3-dimensional vector space Ris described.
2 2 c 2 2 Cin t t Let ƒ:Z→(G)in be a multivector feature map and let w:Z→(G)be a multivector kernel, then [[Lƒ]★w](x)=[L[ƒ★w]](x).
Proof.
2,0 Implementation of Cl(R) layers is now discussed in more detail. One can implement a 2D Clifford CNN layer using Equation 30 where
correspond to 4 different kernels representing one 2D multivector kernel, i.e. 4 different convolution layers, and
correspond to the scalar, vector and bivector parts of the input multivector field. The channels of the different layers represent different stacks of scalars, vectors, and bivectors. All kernels have the same number of input and output channels (number of input and output multivectors), and thus the channels mixing occurs for the different terms of Equations 30, 45 individually. Lastly, usually not all parts of the multivectors are present in the input vector fields. This can easily be accounted for by just omitting the respective parts of Equations 30, 45. A similar reasoning applies to the output vector fields.
0,2 j i,j An alternative parameterization to the Clifford CNN layer introduced in Equation 9 can be realized by using the isomorphism of the Clifford algebra Cl(R) to quaternions. This alternative parameterization takes advantage of the fact that a quaternion rotation can be realized by a matrix multiplication. Using the isomorphism, one can represent the feature maps ƒand filters was quaternions:
j i,j j i,j j i,j −1 j i,j j i,j i,j j i,j i,j Leveraging this quaternion representation, the alternative parameterization of a product between the feature map ƒand wcan be realized. To be more precise, consider a composite operation that results in a scalar quantity and a quaternion rotation, where the latter acts on the vector part of the quaternion ƒand only produces nonzero expansion coefficients for the vector part of the quaternion output. A quaternion rotation wƒ(w)acts on the vector part (î, ĵ, {circumflex over (k)}) of ƒ, and can be algebraically manipulated into a vector-matrix operation Rƒ, where R: H→H is built up from the elements of w. In other words, one can transform the vector part (î, ĵ, {circumflex over (k)}) of ƒ∈H via a rotation matrix Rthat is built from the scalar and vector part (1, (î, ĵ, {circumflex over (k)}) of w∈H. Altogether, a rotational multivector filter
j acts on the feature map ƒthrough a rotational transformation
2 2 c acting on vector and bivector parts of the multivector feature map ƒ: Z→(G)in, and an additional scalar response of the multivector filters:
i,j which is the scalar output of Equation 30. The rotational matrix R(y−x) in written form is:
is the normalized filter with
i,j The dependency (y−x) is omitted inside the rotation matrix Rfor clarity.
0 1 2 12 13 23 123 0 1 2 12 13 23 123 Analogous to the 2-dimensional case, one can implement a 3D Clifford CNN layer using Equation 45, where {b, b, b, b, b, b, b} correspond to 8 different kernels representing one 3D multivector kernel, i.e. 8 different convolution layers, and {a, a, a, a, a, a, a} correspond to the scalar, vector, bivector, and trivector parts of the input multivector field.
Different normalization schemes have been proposed to stabilize and accelerate training DNNs. Their standard formulations apply only to real values. Simply translating and scaling multivectors such that their mean is 0 and their variance is 1 is insufficient because it does not ensure equal variance across all components.
0 1 1 2 2 12 1 2 A batch normalization formulation can be applied to complex values. Build on the same principles can provide an appropriate batch normalization scheme for multivectors. For 2D multivectors of the form a=a+ae+ae+aee, one can formulate the problem of batch normalization as that of whitening 4D vectors:
where the covariance matrix V is
The shift parameter β is a multivector parameter with 4 learnable components and the scaling parameter y is 4×4 positive semi-definite matrix. The multivector batch normalization is defined as:
When the batch sizes are small, it can be more appropriate to use Group Normalization or Layer Normalization. These can be derived with appropriate application of Equation 47 along appropriate tensor dimensions.
Similar to Clifford normalization, quaternion initialization schemes can be adapted to Clifford layers. Effectively, tighter bounds are required for the uniform distribution form which Clifford weights are sampled. However, despite intensive studies no performance gains were observed over default PyTorch initialization schemes for 2-dimensional experiments. However, 3-dimensional implementations necessitate much smaller initialization values (e.g., about a factor of about ⅛).
Clifford convolutions satisfy the property of equivariance under translation of the multivector inputs, as shown. However, the current definition of Clifford convolutions is not equivariant under multivector rotations or reflections. Here, a general kernel constraint is derived which allows building generalized Clifford convolutions which are equivariant with respect to rotations or reflections of the multivectors. That is, equivariance of a Clifford layer under rotations and reflections (i.e. orthogonal transformations) is proved if the multivector kernel multivector filters
satisfies the constraint:
in for 0≤j<c. First define an orthogonal transformation on a multivector by,
2 c out where u and ƒ are multivectors which are multiplied using the geometric product. The minus sign is picked up from reflections but not by rotations, i.e. it depends on the parity of the transformation. This construction is called a “versor” product. The above construction makes it immediately clear that T(ƒg)=(Tƒ)(Tg). Tx means an orthogonal transformation of a Euclidean vector (which can in principle also be defined using versors). To show equivariance, prove for multivectors ƒ: Z→(G)in and a set of cmultivector filters
that:
That is, if the input multivector field transforms as a multivector, and the kernel satisfies the stated equivariance constraint, then the output multivector field also transforms properly as a multivector. Note that T might act differently on the various components (scalars, vectors, pseudoscalars, pseudovectors) under rotations and/or reflections.
Now,
i where in the fourth line variables are transformed as y→y′, in the fifth line the invariance of the summation “measure” under T is used, in the sixth line the transformation property of ƒ and equivariance for wis used, in the seventh line the property of multivectors is used, and in the eighth line linearity of T is used.
2 3 2 3 n 1 n An implementation of Clifford Fourier layers for multivectors in Gand G, i.e. multivectors of Clifford algebras generated by the 2-dimensional vector space Rand the 3-dimensional vector space Ris now derived. In arbitrary dimension n, the Fourier transform {circumflex over (ƒ)}(ξ)={ƒ}(ξ) for a continuous n-dimensional complex-valued signal ƒ(x)=ƒ(x, . . . , x):R→C is defined as:
n★ provided that the integral exists, where x and ξ are n-dimensional vectors and <x, ξ> is the contraction of x and ξ. Usually, <x, ξ> is the inner product, and ξ is an element of the dual vector space R. The inversion theorem states the back-transform from the frequency domain into the spatial domain:
One can rewrite the Fourier transform of Equation 53 in coordinates:
1 n 1 n n The discrete counterpart of Equation 53 transforms an n-dimensional complex signal ƒ(x)=ƒ(x, . . . , x):R→C at M× . . . ×Mgrid points into its complex Fourier modes via:
1 n M1 Mn where (ξ, . . . , ξ)∈Z. . . Z. Fast Fourier transforms (FFTs) immensely accelerate the computation of the transformations of Equation 56 by factorizing the discrete Fourier transform matrix into a product of sparse (mostly zero) factors.
2 2 2 Analogous to Equation 53, the Clifford Fourier transform and the respective inverse transform for multivector valued functions ƒ(x):R→Gand vectors x, ξ∈Rare defined as:
provided that the integrals exist. The differences to Equations 53 and 54 are that ƒ(x) and {circumflex over (ƒ)}(ξ) represent multivector fields in the spatial and the frequency domain, respectively, and that the pseudoscalar i2=e1e2 is used in the exponent. Inserting the definition of multivector fields, one can rewrite Equation 57 as Equation 14.
s/v The discretized versions of the spinor/vector part ({circumflex over (ƒ)}) reads analogously to Equation 56:
1 2 M1 Mn where again (ξ,ξ)∈Z×Z. Similar to Fourier Neural Operators (FNOs) where weight tensors are applied pointwise in the Fourier space, one can apply multivector weight tensors
point-wise. Fourier modes above cut-off frequencies
are set to zero. In doing so, one modifies the Clifford Fourier modes
2,0 via the geometric product. The Clifford Fourier modes follow naturally when combining spinor and vector parts of Equation 59. Analogously to FNOs, higher order modes are cut off. Finally, the residual connections used in FNO layers is replaced by a multivector weight matrix realized as Clifford convolution, ideally a Cl(R) convolution layer.
i2s 2 What follows is proof of the 2D Clifford convolution theorem for multivector valued filters applied from the right, such that filter operations are consistent with Clifford convolution layers. First, it is shown that the Clifford kernel commutes with the spinor and anti-commutes with the vector part of multivectors. Write the product aefor every scalar s∈R and multivector a∈Gas
2 2 1 2 1 1 2 1 2 1 2 1 For the basis of the spinor part, note 1i=i1, and for the basis of the vector part ei=eee=−eee=−ie. Thus, the Fourier kernel
commutes with the spinor part, and anti-commutes with the vector part of a. What follows is a proof of the convolution theorem for the commuting spinor and the anti-commuting vector part of a.
2 2 2 2 2 2 s v s v Let the field ƒ:R→Gbe multivector valued, the filter k:R→Gbe spinor valued, and the filter k:R→Gbe vector valued, and let{ƒ},{k},{k} exist, then
3 3 3 Analogous to Equation 53, the Clifford Fourier transform and the respective inverse transform for multivector valued functions ƒ:R→Gand vectors x, ξ∈Rare defined as:
3 3 provided that the integrals exist. A multivector valued function ƒ:R→G,
3 1 2 3 can be expressed via the pseudoscalar i=eeeas:
0 0 123 3 1 1 23 3 2 2 31 3 3 3 12 3 0 1 2 3 3 4 One can obtain a 3-dimensional Clifford Fourier transform by applying four standard Fourier transforms for the four dual pairs ƒ=ƒ(x)+ƒ(x)i, ƒ=ƒ(x)+ƒ(x)i, ƒ=ƒ(x)+ƒ(x)i, and ƒ=ƒ(x)+ƒ(x)i, which all can be treated as a complex-valued signal ƒ, ƒ, ƒ, ƒ:R→C. Consequently, ƒ(x) can be understood as an element of C. The 3D Clifford Fourier transform is the linear combination of four classical Fourier transforms:
Analogous to the 2-dimensional Clifford Fourier transform, one can apply multivector weight tensors
point-wise. Fourier modes above cut-off frequencies
are set to zero. In doing so, one modifies the Clifford Fourier modes
3,0 via the geometric product. The Clifford Fourier modes follow naturally when combining the four parts of Equation 69. Finally, the residual connections used in FNO layers is replaced by a multivector weight matrix realized as Clifford convolution, such as a Cl(R) convolution layer.
i 3 s 3 First, it is shown that the Clifford kernel commutes with the spinor and anti-commutes with the vector part of multivectors. Write the product aefor every scalar s∈R and multivector a∈Gas
3 In contrast to the 2-dimensional Clifford Fourier transform, now all four parts of the multivector of Equation 68 commute with i:
3 3 3 3 a a Let the field ƒ:R→Gbe multivector valued, the filter k:R→Gbe multivector valued, and let{ƒ},{k} exist, then
3,0 2,0 3,0 One can implement a 2D Clifford Fourier layer by applying two standard Fourier transforms on the dual pairs of Equation 14. These dual pairs can be treated as complex valued inputs. Similarly, one implement a 3D Clifford Fourier layer by applying four standard Fourier transforms on the dual pairs of, for example, ClEquations 40-Equation 43). Since Clifford convolution theorems hold both for the vector and the spinor parts and for the four dual pairs for Cland Cl, respectively, we multiply the modes in the Fourier space using the geometric product. Finally, we apply an inverse Fourier transformation and resemble the multivectors in the spatial domain.
12 FIG. 1200 1200 1220 1222 illustrates, by way of example, a diagram of an embodiment of a methodfor machine learning (ML) modeling of a multivector system that operates on a multivector object. The methodas illustrated includes receiving, by an ML model, the multivector object as an input that represents a state of the multivector system, at operation; and operating, by the ML model and using a Clifford layer that includes neurons that implement a multivector kernel, on the multivector input to generate a multivector output that represents the state of the multivector system responsive to the multivector input, at operation. The multivector input represents a change in state (e.g., stimulus) in the multivector system.
1200 1200 1200 1200 The methodcan further include, wherein the Clifford layer is a Clifford Fourier layer, a Clifford convolution layer, or an equivariant Clifford convolution layer. The methodcan further include, wherein the Clifford layer is a Clifford Fourier Neural Operator (FNO) that includes a Clifford Fourier layer and a Clifford convolution layer. The methodcan further include, wherein the multivector system is an electromagnetic field or a dynamic fluid. The methodcan further include, wherein the Clifford layer is a Clifford convolution and the method further comprises constraining a kernel for a Clifford convolution layer such that the Clifford convolution layer is equivariant.
1200 1200 1200 1200 The methodcan further include, wherein the multivector object includes two or more of a scalar, a vector, a bi-vector, or a tri-vector. The methodcan further include, wherein the multivector object includes the vector. The methodcan further include, wherein the multivector object includes the bi-vector. The methodcan further include, wherein the multivector object includes the tri-vector.
102 108 Artificial Intelligence (AI) is a field concerned with developing decision-making systems to perform cognitive tasks that have traditionally required a living actor, such as a person. Neural networks (NNs) are computational structures that are loosely modeled on biological neurons. Generally, NNs encode information (e.g., data or decision making) via weighted connections (e.g., synapses) between nodes (e.g., neurons). Modern NNs are foundational to many AI applications, such as object or condition recognition, device behavior modeling or the like. The ML modelor a portion thereof, such as the Clifford neural layer, can include or be implemented using one or more NNs.
Many NNs are represented as matrices of weights (sometimes called parameters) that correspond to the modeled connections. NNs operate by accepting data into a set of input neurons that often have many outgoing connections to other neurons. At each traversal between neurons, the corresponding weight modifies the input and is tested against a threshold at the destination neuron. If the weighted value exceeds the threshold, the value is again weighted, or transformed through a nonlinear function, and transmitted to another neuron further down the NN graph—if the threshold is not exceeded then, generally, the value is not transmitted to a down-graph neuron and the synaptic connection remains inactive. The process of weighting and testing continues until an output neuron is reached; the pattern and values of the output neurons constituting the result of the NN processing.
The optimal operation of most NNs relies on accurate weights. However, NN designers do not generally know which weights will work for a given application. NN designers typically choose a number of neuron layers or specific connections between layers including circular connections. A training process may be used to determine appropriate weights by selecting initial weights.
In some examples, initial weights may be randomly selected. Training data is fed into the NN, and results are compared to an objective function that provides an indication of error. The error indication is a measure of how wrong the NN's result is compared to an expected result. This error is then used to correct the weights. Over many iterations, the weights will collectively converge to encode the operational data into the NN. This process may be called an optimization of the objective function (e.g., a cost or loss function), whereby the cost or loss is minimized.
A gradient descent technique is often used to perform objective function optimization. A gradient (e.g., partial derivative) is computed with respect to layer parameters (e.g., aspects of the weight) to provide a direction, and possibly a degree, of correction, but does not result in a single correction to set the weight to a “correct” value. That is, via several iterations, the weight will move towards the “correct,” or operationally useful, value. In some implementations, the amount, or step size, of movement is fixed (e.g., the same from iteration to iteration). Small step sizes tend to take a long time to converge, whereas large step sizes may oscillate around the correct value or exhibit other undesirable behavior. Variable step sizes may be attempted to provide faster convergence without the downsides of large step sizes.
Backpropagation is a technique whereby training data is fed forward through the NN—here “forward” means that the data starts at the input neurons and follows the directed graph of neuron connections until the output neurons are reached—and the objective function is applied backwards through the NN to correct the synapse weights. At each step in the backpropagation process, the result of the previous step is used to correct a weight. Thus, the result of the output neuron correction is applied to a neuron that connects to the output neuron, and so forth until the input neurons are reached. Backpropagation has become a popular technique to train a variety of NNs. Any well-known optimization algorithm for back propagation may be used, such as stochastic gradient descent (SGD), Adam, etc.
13 FIG. 13 FIG. 1305 1310 1310 1305 1307 1310 1305 is a block diagram of an example of an environment including a system for neural network training. One or more of Clifford neural layers and other parts of an ML model discussed herein can be at least partially implemented using the system of. The system includes an artificial NN (ANN)that is trained using a processing node. The processing nodemay be a central processing unit (CPU), graphics processing unit (GPU), field programmable gate array (FPGA), digital signal processor (DSP), application specific integrated circuit (ASIC), or other processing circuitry. In an example, multiple processing nodes may be employed to train different layers of the ANN, or even different nodeswithin layers. Thus, a set of processing nodesis arranged to perform the training of the ANN.
1310 1315 1305 1305 1307 1307 1308 1315 1305 The set of processing nodesis arranged to receive a training setfor the ANN. The ANNcomprises a set of nodesarranged in layers (illustrated as rows of nodes) and a set of inter-node weights(e.g., parameters) between nodes in the set of nodes. In an example, the training setis a subset of a complete training set. Here, the subset may enable processing nodes with limited storage resources to participate in training the ANN.
1317 1305 1307 1305 The training data may include multiple numerical values representative of a domain, such as an image feature, or the like. Each value of the training or inputto be classified after ANNis trained, is provided to a corresponding nodein the first layer or input layer of ANN. The values propagate through the layers and are changed by the objective function.
1320 1317 1307 1305 1305 1305 1307 As noted, the set of processing nodes is arranged to train the neural network to create a trained neural network. After the ANN is trained, data input into the ANN will produce valid classifications(e.g., the input datawill be assigned into categories), for example. The training performed by the set of processing nodesis iterative. In an example, each iteration of the training the ANNis performed independently between layers of the ANN. Thus, two distinct layers may be processed in parallel by different members of the set of processing nodes. In an example, different layers of the ANNare trained on different hardware. The members of different members of the set of processing nodes may be located in different packages, housings, computers, cloud-based resources, etc. In an example, each iteration of the training is performed independently between nodes in the set of nodes. This example is an additional parallelization whereby individual nodes(e.g., neurons) are trained independently. In an example, the nodes are trained on different hardware.
14 FIG. 13 FIG. 14 FIG. 1400 102 108 600 700 1200 1400 1400 1402 1403 1410 1412 1400 1400 illustrates, by way of example, a block diagram of an embodiment of a machine(e.g., a computer system) to implement one or more embodiments. One or more of the ML model, Clifford neural layer, Clifford convolution layer, Clifford FNO layer, the method, the training system of, or a component or operations thereof can be implemented, at least in part, using a component of the machine. One example machine(in the form of a computer), may include a processing unit, memory, removable storage, and non-removable storage. Although the example computing device is illustrated and described as machine, the computing device may be in different forms in different embodiments. For example, the computing device may instead be a smartphone, a tablet, smartwatch, or other computing device including the same or similar elements as illustrated and described regarding. Further, although the various data storage elements are illustrated as part of the machine, the storage may also or alternatively include cloud-based storage accessible via a network, such as the Internet.
1403 1414 1408 1400 1414 1408 1410 1412 Memorymay include volatile memoryand non-volatile memory. The machinemay include—or have access to a computing environment that includes—a variety of computer-readable media, such as volatile memoryand non-volatile memory, removable storageand non-removable storage. Computer storage includes random access memory (RAM), read only memory (ROM), erasable programmable read-only memory (EPROM) & electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD ROM), Digital Versatile Disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices capable of storing computer-readable instructions for execution to perform functions described herein.
1400 1406 1404 1416 1404 1406 1400 The machinemay include or have access to a computing environment that includes input, output, and a communication connection. Outputmay include a display device, such as a touchscreen, that also may serve as an input component. The inputmay include one or more of a touchscreen, touchpad, mouse, keyboard, camera, one or more device-specific buttons, one or more sensors integrated within or coupled via wired or wireless data connections to the machine, and other input components. The computer may operate in a networked environment using a communication connection to connect to one or more remote computers, such as database servers, including cloud-based servers and storage. The remote computer may include a personal computer (PC), server, router, network PC, a peer device or other common network node, or the like. The communication connection may include a Local Area Network (LAN), a Wide Area Network (WAN), cellular, Institute of Electrical and Electronics Engineers (IEEE) 802.11 (Wi-Fi), Bluetooth, or other networks.
1402 1400 1418 1402 Computer-readable instructions stored on a computer-readable storage device are executable by the processing unit(sometimes called processing circuitry) of the machine. A hard drive, CD-ROM, and RANI are some examples of articles including a non-transitory computer-readable medium such as a storage device. For example, a computer programmay be used to cause processing unitto perform one or more methods or algorithms described herein.
The operations, functions, or algorithms described herein may be implemented in software in some embodiments. The software may include computer executable instructions stored on computer or other machine-readable media or storage device, such as one or more non-transitory memories (e.g., a non-transitory machine-readable medium) or other type of hardware-based storage devices, either local or networked. Further, such functions may correspond to subsystems, which may be software, hardware, firmware, or a combination thereof. Multiple functions may be performed in one or more subsystems as desired, and the embodiments described are merely examples. The software may be executed on processing circuitry, such as can include a digital signal processor, ASIC, microprocessor, central processing unit (CPU), graphics processing unit (GPU), field programmable gate array (FPGA), or other type of processor operating on a computer system, such as a personal computer, server, or other computer system, turning such computer system into a specifically programmed machine. The processing circuitry can, additionally or alternatively, include electric and/or electronic components (e.g., one or more transistors, resistors, capacitors, inductors, amplifiers, modulators, demodulators, antennas, radios, regulators, diodes, oscillators, multiplexers, logic gates, buffers, caches, memories, GPUs, CPUs, field programmable gate arrays (FPGAs), or the like). The terms computer-readable medium, machine readable medium, and storage device do not include carrier waves or signals to the extent carrier waves and signals are deemed too transitory.
Example 1 includes a method for machine learning (ML) modeling of a multivector system that operates on a multivector object, the method comprising receiving, by an ML model, the multivector object as an input that represents a state of the multivector system, and operating, by the ML model and using a Clifford layer that includes neurons that implement a multivector kernel, on the multivector input to generate a multivector output that represents the state of the multivector system responsive to the multivector input.
In Example 2, Example 1 further includes, wherein the Clifford layer is a Clifford Fourier layer, a Clifford convolution layer, or an equivariant Clifford convolution layer.
In Example 3, at least one of Examples 1-2 further includes, wherein the Clifford layer is a Clifford Fourier Neural Operator (FNO) that includes a Clifford Fourier layer and a Clifford convolution layer.
In Example 4, at least one of Examples 1-3 further includes, wherein the multivector system is an electromagnetic field or a dynamic fluid.
In Example 5, at least one of Examples 1-4 further includes, wherein the multivector object includes two or more of a scalar, a vector, a bi-vector, or a tri-vector.
In Example 6, Example 5 further includes, wherein the multivector object includes the vector.
In Example 7, at least one of Examples 5-6 further includes, wherein the multivector object includes the bi-vector.
In Example 8, at least one of Examples 5-7 further includes, wherein the multivector object includes the tri-vector.
In Example 9, at least one of Examples 1-8 further includes, wherein the Clifford layer is a Clifford convolution and the method further comprises constraining a kernel for a Clifford convolution layer such that the Clifford convolution layer is equivariant.
Example 10 includes a system comprising processing circuitry and one or more memories including parameters for neurons trained for machine learning (ML) modeling of a multivector system that operates on a multivector object and instructions that, when executed by the processing circuitry, cause the processing circuitry to perform operations for implementing an ML model using the method of one of Examples 1-9.
1 9 Example 11 includes a computer-readable medium including instructions that, when executed by a machine, cause the machine to perform the method of one of claims-.
Although a few embodiments have been described in detail above, other modifications are possible. For example, the logic flows depicted in the figures do not require the order shown, or sequential order, to achieve desirable results. The desirable for embodiments can include the user having confidence in the state of their data, settings, controls, and secrets before, during, and after a migration to a new version of an application. Using multiple factors to check data state, integrity, presence, and absence before and after the migration can increase confidence. Other steps may be provided, or steps may be eliminated, from the described flows, and other components may be added to, or removed from, the described systems. Other embodiments may be within the scope of the following claims.
Cooperative Patent Classification codes for this invention. Click any code to explore related patents in that topic.
December 22, 2022
July 14, 2026
Browse 5M+ US patents with plain-English claim translations and AI-generated analysis.