SXEAI: Reverse-Engineering the Mathematics Learned by a Neural Sextic Solver

Abstract

Can a neural network learn not only to solve an algebraic problem, but also to develop an internal mathematical representation of the problem that we can reverse-engineer?

This is the main question behind SXEAI, an experimental project focused on neural networks that recover the six roots of a general polynomial of degree six from its coefficients.

The objective is deliberately different from a conventional machine-learning benchmark. We are not primarily interested in obtaining the smallest possible root prediction error. Instead, we want to understand what the network computes internally.

Across a sequence of experiments, SXEAI repeatedly developed hidden representations that could be approximated by compact combinations of symmetric polynomial quantities, power sums, centered moments, depressed-polynomial coefficients, and related algebraic features. Some of these internal coordinates could be discovered blindly, reproduced on unseen data, and even replaced causally by their symbolic approximations without substantially changing downstream network behavior.

The resulting picture is not yet a discovered closed-form solution of the general sextic equation. It is something more precise and, in our view, more interesting: evidence that a neural algebraic solver can spontaneously organize its hidden state into a structured mathematical coordinate system that can be partially reconstructed and experimentally manipulated.


1. The question behind SXEAI

A standard neural network experiment asks a simple question:

Can the network learn the mapping from inputs to outputs?

For a sextic polynomial,

P(z)=z6+p5z5+p4z4+p3z3+p2z2+p1z+p0, P(z)=z^6+p_5z^5+p_4z^4+p_3z^3+p_2z^2+p_1z+p_0,

the conventional task would be:

(p5,p4,p3,p2,p1,p0)⟢(r1,…,r6), (p_5,p_4,p_3,p_2,p_1,p_0) \longrightarrow (r_1,\ldots,r_6),

where

P(z)=∏i=16(zβˆ’ri). P(z)=\prod_{i=1}^{6}(z-r_i).

One can then measure MAE, RMSE, maximum root error, or polynomial reconstruction error.

SXEAI asks a different question:

What mathematical representation does the network build internally? \boxed{\text{What mathematical representation does the network build internally?}}

The network may not need to directly represent each root at every stage. It could instead construct intermediate quantities such as

βˆ’p56, -\frac{p_5}{6},

power sums

sk=βˆ‘irik, s_k=\sum_i r_i^k,

centered moments

qk=16βˆ‘i(riβˆ’ΞΌ)k, q_k=\frac16\sum_i(r_i-\mu)^k,

or more complicated combinations such as

s2s3,s2s4,q3q4. s_2s_3,\qquad s_2s_4,\qquad q_3q_4.

The central idea of the project is therefore to treat the network itself as a mathematical object that can be reverse-engineered.


2. Why the sextic problem?

The sextic is an interesting test case because it is already beyond the regime where one should expect a simple universal elementary formula for the roots.

For low-degree equations, classical symbolic methods provide familiar closed forms. At degree six, the algebraic structure becomes substantially richer.

This makes the sextic a useful environment for asking:

Does a neural solver rediscover recognizable algebraic coordinates, or does it develop a completely different internal representation?

We do not assume in advance that the answer must involve radicals.

That distinction became increasingly important as the project progressed.


3. The model

SXEAI uses a custom layer called SexticLayer.

A layer first computes

z=Wx+b, z=Wx+b,

then applies a trainable sixth-degree polynomial

Az6+Bz5+Cz4+Dz3+Ez2+Fz+G, Az^6+Bz^5+Cz^4+Dz^3+Ez^2+Fz+G,

followed by normalization and GELU.

The polynomial activation is initialized close to a quadratic form, so the network starts from a relatively simple nonlinear regime but can learn higher-order behavior.

The main complex model used in the later experiments has approximately 1.2 million parameters and consists of 24 such layers with hidden width 223.

The final network maps a 12-dimensional real representation of the polynomial coefficients to 12 real values representing six complex roots:

[β„œr1,β„‘r1,…,β„œr6,β„‘r6]. [\Re r_1,\Im r_1,\ldots,\Re r_6,\Im r_6].

This architecture is intentionally unusual. Rather than hiding all nonlinearity inside a standard ReLU or GELU network, we give the model explicit trainable polynomial operations and then ask what mathematical structure emerges.


4. Data generation

The polynomial is generated from its roots.

For six roots

r1,…,r6, r_1,\ldots,r_6,

we construct

P(z)=∏i=16(zβˆ’ri). P(z)=\prod_{i=1}^{6}(z-r_i).

The network does not receive the roots. It receives only the resulting coefficients.

For the complex experiments, roots are represented by real and imaginary parts:

ri=ai+ibi. r_i=a_i+ib_i.

The root generation, normalization, coefficient construction, ordering and scaling are kept fixed across the main comparisons.

The training protocol used in the main experiments includes:

  • 350,000 training samples;
  • batch size 768;
  • 125 epochs;
  • AdamW;
  • learning rate (3\times10^{-4});
  • weight decay (10^{-4});
  • gradient clipping at 1.0;
  • SmoothL1 loss.

The dataset mixes real, natural/integer-like and fully complex root configurations.


5. The first discovery: the network finds the center

One of the earliest and most persistent observations was the appearance of

ΞΌ=βˆ’p56. \boxed{\mu=-\frac{p_5}{6}}.

Since

p5=βˆ’βˆ‘iri, p_5=-\sum_i r_i,

we have

ΞΌ=16βˆ‘iri. \mu=\frac16\sum_i r_i.

In other words, the first layers repeatedly discovered the mean of the roots.

This was not a one-off correlation.

Across different architectures and training runs, early hidden units consistently exhibited strong alignment with this quantity.

That suggested an important possibility:

the network may first transform the problem into a coordinate system centered around the roots before attempting to recover their detailed structure.


6. From coefficients to symmetric quantities

The next step was to search hidden activations against a library of algebraically meaningful quantities.

Among the most important were the power sums

sk=βˆ‘irik, s_k=\sum_i r_i^k,

together with centered moments

qk=16βˆ‘i(riβˆ’ΞΌ)k q_k=\frac16\sum_i(r_i-\mu)^k

and the coefficients of the depressed sextic obtained after translating the variable by the root mean.

Newton identities provide the connection between these representations.

For example,

s1=βˆ‘iri=βˆ’p5, s_1=\sum_i r_i=-p_5,

and

s2,β€…β€Šs3,… s_2,\;s_3,\ldots

can be expressed recursively in terms of the polynomial coefficients.

This means the network has access, through ordinary algebraic transformations, to several equivalent coordinate systems describing the same root configuration.

The symbolic experiments suggested that it was not merely using raw coefficient coordinates.

It repeatedly rediscovered combinations of the more natural symmetric quantities.


7. Deep layers developed richer combinations

In deeper networks we began to find hidden units whose activations could be approximated by combinations such as

as3+bs5+c s2s3, a s_3+b s_5+c\,s_2s_3,

or

as6+b s2s4+c q4. a s_6+b\,s_2s_4+c\,q_4.

Examples from earlier high-performing real-root models included expressions such as

1.7637s3+0.2457s5βˆ’1.2059s2s3 1.7637s_3+0.2457s_5-1.2059s_2s_3

with (R^2) around 0.80, and

βˆ’0.7720s1βˆ’0.7597s3+0.6254s2s3 -0.7720s_1-0.7597s_3+0.6254s_2s_3

with (R^2) around 0.89.

The exact coefficients are not the scientific result by themselves.

The important observation is the repeated vocabulary:

s1,s2,s3,s5,s6,s2s3,s2s4,… s_1,\quad s_2,\quad s_3,\quad s_5,\quad s_6, \quad s_2s_3,\quad s_2s_4,\ldots

These objects have a clear mathematical interpretation as symmetric functions of the roots.


8. The architecture did not scale monotonically

One of the strongest negative results appeared when we increased network depth.

Larger models did not automatically produce better numerical or symbolic behavior.

Some 32-layer configurations entered a collapsed regime in which the output became almost constant across very different polynomial inputs.

A representative failure had approximately

MAEβ‰ˆ12.5, \mathrm{MAE}\approx12.5,

compared with values near 1 or below 1 for successful real-root runs.

This happened even when the parameter budget was large.

Later, a particularly clean comparison used approximately the same parameter count:

24L/600K 24L/600K

versus

32L/600K. 32L/600K.

The 24-layer model worked normally, while the 32-layer model collapsed after only a few epochs.

This is an important result because it shows:

more depth does not automatically produce more algebraic reasoning \boxed{\text{more depth does not automatically produce more algebraic reasoning}}

and, in this architecture, depth interacts strongly with optimization stability.


9. A second important observation: optimization is part of the phenomenon

Some models trained well for many epochs and then suddenly deteriorated.

For example, a 16-layer model reached a very low training loss before a later catastrophic transition.

Other models collapsed much earlier.

This means that SXEAI is not just about the final trained state.

The trajectory of learning itself contains information.

A model can pass through a phase where its internal representation is highly structured, and later destroy that representation through optimization instability.

That led to one of the most interesting changes in the project:

Instead of analysing only the final checkpoint, we began analysing the hidden representation as a function of training epoch.


10. A stricter symbolic protocol

Our early symbolic experiments had several weaknesses.

First, some features were algebraically dependent.

For example,

p5=βˆ’6ΞΌ. p_5=-6\mu.

Therefore

p52=36ΞΌ2. p_5^2=36\mu^2.

Treating these as independent regression coordinates can create misleading sparse formulas.

Second, fitting and evaluating on the same probe points can overestimate symbolic quality.

Third, a very high (R^2) can sometimes be produced by a nearly constant hidden unit.

We therefore moved toward a stricter procedure:

  • separate symbolic fit and holdout sets;
  • larger feature libraries;
  • numerical rank reduction;
  • removal of obvious aliases;
  • sparse formulas with one to three main terms;
  • explicit checks for hidden activation variance and non-triviality.

This changed the interpretation of our symbolic results substantially.


11. The complex solver

The later SXEAI experiments extended the problem from six real roots to six complex roots.

Each root is

ri=ai+ibi, r_i=a_i+ib_i,

so the problem becomes a 12-dimensional real input/output mapping.

The approximately 600K-parameter model was able to recover the complex roots with non-trivial accuracy, while keeping real and natural root configurations easier than general complex configurations.

A stronger 600K complex model produced a 100K independent-test result around:

MAEβ‰ˆ0.074, \mathrm{MAE}\approx0.074,

with substantially larger errors on fully complex examples than on real or natural ones.

The important point here is not that SXEAI has solved the general complex sextic to high precision.

It has not.

The important point is that a relatively small neural architecture can learn a useful approximation to the inverse coefficient-to-root map even when the roots are genuinely complex.


12. The 1.2M model and an unexpected result

We then approximately doubled the parameter budget.

The larger 24-layer model had:

1,202,651 1,202,651

parameters.

Its training trajectory was highly unstable.

The best validation checkpoint occurred around epoch 2. After that, the network entered a poorer regime and remained there.

This created an unusual opportunity.

Instead of looking only at the final model, we could study the early checkpoint before collapse.

That checkpoint turned out to contain extremely strong symbolic structure.


13. Blind symbolic discovery

One of the most important controls was a blind observer.

The observer did not receive the formulas we had previously discovered.

It only saw:

input coefficientsβ†’hidden activation. \text{input coefficients} \rightarrow \text{hidden activation}.

The observer reconstructed power sums and other algebraic features and independently searched for sparse explanations.

It rediscovered hidden coordinates with extremely high holdout (R^2), in some cases around

R2β‰ˆ0.9997βˆ’0.9999. R^2\approx0.9997-0.9999.

For example, one hidden coordinate could be described approximately by

0.046022βˆ’0.007446p1+0.001555p12βˆ’0.000385p2βˆ’0.000121p13. 0.046022 -0.007446p_1 +0.001555p_1^2 -0.000385p_2 -0.000121p_1^3.

The scientific significance is not the numerical coefficients.

The important point is that an observer that did not know the earlier answers independently converged to the same algebraic vocabulary.


14. A shared symbolic vocabulary emerged

Across many hidden units, the same small family of features repeatedly appeared:

p1,p12,p2,p13,p1p2,p3, p_1,\quad p_1^2,\quad p_2,\quad p_1^3,\quad p_1p_2,\quad p_3,

which, through Newton identities, corresponds to quantities involving the root center and centered moments.

Using

p1=6ΞΌ, p_1=6\mu,

p2=6(ΞΌ2+q2), p_2=6(\mu^2+q_2),

and

p3=6(ΞΌ3+3ΞΌq2+q3), p_3=6(\mu^3+3\mu q_2+q_3),

the apparently different symbolic formulas can often be interpreted as a common coordinate system based on

ΞΌ,q2,q3,… \mu,q_2,q_3,\ldots

This was one of the strongest conceptual results of the project.

The network does not appear to discover completely unrelated formulas for every hidden unit.

Instead, many hidden units seem to speak a shared algebraic language.


15. A concrete internal coordinate

One particularly strong example was a hidden coordinate of the approximate form

hβ‰ˆ0.062438ΞΌ3βˆ’0.025100ΞΌ2q2βˆ’0.241301ΞΌ+0.018231q2+0.030144. h \approx 0.062438\mu^3 -0.025100\mu^2q_2 -0.241301\mu +0.018231q_2 +0.030144.

On fresh data this type of representation reached

R2β‰ˆ0.9967. R^2\approx0.9967.

Another coordinate had a structure approximately like

hβ‰ˆ0.768039ΞΌ2βˆ’0.036781ΞΌq3βˆ’2.116605ΞΌ+0.059717q2+1.379858. h \approx 0.768039\mu^2 -0.036781\mu q_3 -2.116605\mu +0.059717q_2 +1.379858.

Its fit was similarly strong.

Again, these are not claims about the final roots.

They are examples of internal coordinates learned by the network.


16. The strongest test: causal substitution

Correlation is not causality.

A symbolic formula can fit a hidden neuron extremely well and still be irrelevant to the actual computation.

To test this, we performed an intervention.

For a selected hidden unit, instead of using

hnetwork, h_{\text{network}},

we replaced it with its symbolic approximation:

hnetwork⟢hsymbolic. h_{\text{network}} \longrightarrow h_{\text{symbolic}}.

We then ran the remainder of the neural network normally.

For several tested hidden coordinates, the downstream prediction barely changed.

Representative output drift values were on the order of

10βˆ’5. 10^{-5}.

This is much stronger evidence than an ordinary symbolic correlation.

It indicates that the symbolic surrogate is not merely visually similar to the hidden neuron.

It is functionally interchangeable with the network's original computation, at least locally and for the tested interventions.


17. Multi-neuron intervention

We extended the experiment by simultaneously replacing several independently discovered symbolic coordinates.

The resulting downstream drift was approximately

5.7Γ—10βˆ’4, 5.7\times10^{-4},

while a matched mean-ablation control produced a much larger drift, roughly

1.6Γ—10βˆ’2. 1.6\times10^{-2}.

This is one of the most important experiments in the project.

It suggests that multiple hidden computations can be replaced by explicit symbolic programs without substantially destroying the network's behavior.

That moves SXEAI beyond:

Β«this neuron correlates with an equationΒ»

toward:

Β«this equation is a viable executable substitute for part of the neural computation.Β»


18. The strange behaviour of late training

The final collapsed checkpoint created another methodological lesson.

Some late hidden units produced symbolic regression scores extremely close to

R2=1. R^2=1.

At first this looked spectacular.

It was not.

The hidden activations had become nearly constant.

A constant or almost constant signal can be approximated almost perfectly by a trivial constant formula.

Therefore:

R2=1β‰ automatic mathematical discovery \boxed{R^2=1\neq\text{automatic mathematical discovery}}

A meaningful symbolic score must be accompanied by:

  • hidden-state variance;
  • coefficient magnitude;
  • complexity;
  • non-trivial feature contribution;
  • independent holdout;
  • and, ideally, a causal intervention.

This lesson is now part of the SXEAI methodology.


19. What about radicals?

At one stage we considered whether the network might be rediscovering a classical radical solution.

We deliberately became more cautious about this interpretation.

For a general sextic, a simple universal radical expression is not the natural expectation.

Therefore there is no reason to assume in advance that the network should internally resemble a hand-derived formula using only nested radicals.

The evidence currently supports a weaker but more robust claim:

the network develops structured algebraic coordinates.

Exactly how those coordinates are assembled into the final roots is still an open question.


20. What about theta functions and Abelian structures?

This became one of the most interesting theoretical possibilities.

The network is not restricted to our symbolic library.

Our library mostly contains polynomial, symmetric, centered and related algebraic quantities.

It does not directly contain:

  • theta functions;
  • Abel-Jacobi coordinates;
  • period matrices;
  • Abelian integrals;
  • inverse Abel maps;
  • more general quasi-periodic special functions.

Therefore, failure to find such functions in our current symbolic search would not prove that the network is not using them.

At the same time, we do not have evidence that the network has discovered a theta-function representation.

That remains a hypothesis.

The correct scientific stance is:

known algebraic structure is demonstrated; deeper analytic structure remains open \boxed{\text{known algebraic structure is demonstrated; deeper analytic structure remains open}}


21. What has actually been demonstrated

At this stage, the strongest defensible conclusions are:

1. Neural sextic solving is possible

A relatively compact polynomial network can learn a useful approximation to the coefficient-to-root map, including genuinely complex roots.

2. The hidden representation is structured

Hidden coordinates repeatedly correlate with natural symmetric quantities such as

ΞΌ,β€…β€Šsk,β€…β€Šqk, \mu,\;s_k,\;q_k,

and combinations of them.

3. The structure is not only a fitting artifact

Independent probe sets and blind symbolic discovery reproduce the same types of algebraic coordinates.

4. The discovered structure can be causally tested

Replacing selected hidden computations with their symbolic approximations causes only very small downstream changes in the tested experiments.

5. Optimization determines whether the representation survives

Increasing depth or width does not guarantee a better solver. Several larger networks entered collapse regimes.

6. Numerical quality and symbolic structure are different axes

A numerically inferior model can contain highly interpretable internal coordinates, while a numerically strong model does not automatically yield a more interpretable representation.


22. What has NOT been demonstrated

We have not shown that SXEAI has:

  • discovered a universal closed-form sextic solution;
  • rediscovered a known resolvent construction;
  • discovered a radical formula;
  • discovered theta functions;
  • discovered an Abelian uniformization;
  • or uncovered a theorem-level new solution of the general sextic equation.

Those claims would go beyond the evidence.

The current result is more specific:

SXEAI spontaneously develops compact, reproducible and causally usable algebraic representations. \boxed{ \text{SXEAI spontaneously develops compact, reproducible and causally usable algebraic representations.} }


23. The emerging picture

The current hypothesis can be summarized schematically as

coefficientsβ†’centeringβ†’symmetric statisticsβ†’nonlinear algebraic mixingβ†’latent mathematical coordinatesβ†’roots \boxed{ \text{coefficients} \rightarrow \text{centering} \rightarrow \text{symmetric statistics} \rightarrow \text{nonlinear algebraic mixing} \rightarrow \text{latent mathematical coordinates} \rightarrow \text{roots} }

The earliest layers often expose simple quantities such as

βˆ’p56, -\frac{p_5}{6},

while deeper layers mix higher-order symmetric information.

The network therefore appears to construct an internal coordinate system rather than directly memorizing six independent output values.

That is the central idea of SXEAI.


24. Why this matters beyond sextics

The project suggests a general methodology for interpretable machine learning.

Instead of asking only:

What does this neuron correlate with?

we can ask a stronger sequence of questions:

observeβ†’symbolically reconstructβ†’test on unseen dataβ†’blindly rediscoverβ†’causally substitute \boxed{ \text{observe} \rightarrow \text{symbolically reconstruct} \rightarrow \text{test on unseen data} \rightarrow \text{blindly rediscover} \rightarrow \text{causally substitute} }

This can potentially be applied to other domains where a neural network learns a hidden mathematical algorithm.

SXEAI therefore becomes less about one polynomial and more about a methodology for reverse-engineering learned computation into executable mathematical programs.


25. Reproducibility

The project is designed so that the model can be released together with its weights and inference code.

The important reproducibility principle is:

The exact checkpoint, architecture definition, preprocessing and coefficient scaling must remain synchronized.

In particular, the released checkpoint must be loaded with the same SexticLayer implementation and the same input normalization used during training.

The following example shows the intended inference workflow.

import torch
from pathlib import Path


DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
WEIGHTS = Path("sxeai_weights.pth")


# ------------------------------------------------------------
# IMPORTANT:
# Use the exact SexticLayer / SXEAI model implementation
# that was used to create the released checkpoint.
#
# The checkpoint is a state_dict, so architecture, parameter
# names and preprocessing must match exactly.
# ------------------------------------------------------------

from sxeai_model import SXEAIComplexSextic


def load_model(weights_path: str):
    model = SXEAIComplexSextic(
        in_dim=12,
        hidden_dim=223,
        depth=24,
        out_dim=12,
    )

    checkpoint = torch.load(
        weights_path,
        map_location=DEVICE,
    )

    # Support either a raw state_dict or a checkpoint dictionary.
    if isinstance(checkpoint, dict) and "state_dict" in checkpoint:
        state_dict = checkpoint["state_dict"]
    else:
        state_dict = checkpoint

    model.load_state_dict(state_dict, strict=True)
    model.to(DEVICE)
    model.eval()

    return model


@torch.no_grad()
def predict(model, x):
    """
    x:
        Tensor of shape [batch, 12].

        The 12 values must use the exact preprocessing
        convention from the training code:

        [Re(p5), Im(p5),
         Re(p4), Im(p4),
         ...
         Re(p0), Im(p0)]
    """

    x = x.to(DEVICE)
    y = model(x)

    return y


def unpack_roots(y):
    """
    Convert [Re(r1), Im(r1), ..., Re(r6), Im(r6)]
    into a complex tensor of shape [batch, 6].
    """

    re = y[:, 0::2]
    im = y[:, 1::2]

    return torch.complex(re, im)


if __name__ == "__main__":
    model = load_model(str(WEIGHTS))

    # Example batch.
    # Replace this tensor with properly preprocessed
    # coefficient vectors from your own sextic polynomials.
    x = torch.randn(4, 12, device=DEVICE)

    y = predict(model, x)
    roots = unpack_roots(y)

    print("Predicted roots:")
    print(roots)

For a public release, the repository should contain:

SXEAI/
β”œβ”€β”€ sxeai_model.py
β”œβ”€β”€ inference.py
β”œβ”€β”€ preprocessing.py
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ README.md
└── weights/
    └── sxeai_weights.pth

The most important file is the exact sxeai_model.py used to generate the checkpoint. This prevents a common reproducibility failure in which the weights are published but the implementation differs slightly from the training model.


26. Recommended public inference interface

The cleanest public interface is:

model = load_model("sxeai_weights.pth")

roots = predict(model, coefficients)

where coefficients already follow the official preprocessing convention.

The output is six complex values:

(r^1,…,r^6). (\hat r_1,\ldots,\hat r_6).

Because roots are intrinsically unordered, downstream applications should preferably evaluate the result using a set-based or optimal matching metric rather than assuming that output position (k) always corresponds to a particular physical root.


27. Final perspective

SXEAI started as an experiment in training a neural network to recover roots of a difficult polynomial.

It became something different.

The interesting result is not simply that a neural network can approximate

coefficients→roots. \text{coefficients}\rightarrow\text{roots}.

The interesting result is that, inside the network, we can observe the emergence of compact mathematical coordinates, reconstruct them symbolically, rediscover them without providing the answers to the search procedure, and experimentally replace parts of the learned computation with explicit formulas.

The project therefore suggests a broader possibility:

A learned neural algorithm can sometimes be reverse-engineered into a mathematical program. \boxed{ \text{A learned neural algorithm can sometimes be reverse-engineered into a mathematical program.} }

SXEAI is an early experiment in that direction.

It does not claim to have solved the general sextic equation in closed form.

It claims something more experimentally grounded:

the neural solver develops structured internal algebra, and that structure can be measured, reconstructed and intervened upon. \boxed{ \text{the neural solver develops structured internal algebra, and that structure can be measured, reconstructed and intervened upon.} }

The next step is to determine how far that reconstruction can go β€” from local hidden coordinates to a complete mathematical description of the computation that maps a sextic polynomial to its roots.


Links to SEAI and CEAI : https://huggingface.co/ALEXFLR/SEAI and https://huggingface.co/ALEXFLR/CEAI.

By ALEX_FLR

You can download weights in another file.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support