Shayan Hajhashemi

Research note · Draft · Updated September 2026

Weak forms, weighted windows, and Lorenz

Can an integrated equation make learned dynamics more robust to noisy observations? In this Lorenz63 benchmark, weak-form SINDy gives up some clean-data efficiency but degrades much more slowly as training noise rises. Giving less weight to windows near the ends of a trajectory does not add a consistent benefit.

The newer experiments sharpen that result rather than replacing it. Known-form integral matching remains strong, direct foundation-model forecasts are useful but belong to a different information track, and a weakly trained Panda-to-SINDy conditioner improves on unseen source systems but fails a reserved Lorenz gate.

Three forecast-survival panels at 0.1%, 1%, and 5% benchmark training noise. Solid curves show state-only learned dynamics, dashed curves include known-form integral and automatic-differentiation Lorenz fits, dotted curves show known-physics surrogates, and dash-dot curves show fixed clean-prefix Panda and Chronos-T5 forecasts. A compact boxed legend below the panels groups all 16 methods into four information-track columns, each headed by its line style.
Forecast survival through five Lyapunov times. Each curve is the fraction of trajectories whose error has remained below the threshold. Solid, dashed, dotted, and dash-dot lines distinguish state-only learning, known equation form with hidden parameters, known-physics forward surrogates, and externally pretrained sequence forecasting. Panda and Chronos receive a fixed 512-sample clean test prefix and are repeated across panels because they do not train on the panel's noisy split. Here “zero-shot” means no target-time optimization; it does not certify that Lorenz was excluded from pretraining. Open full-size figure ↗ · Download PDF

The benchmark

The reference system is Lorenz63, with \(\sigma = 10\), \(\rho = 28\), and \(\beta = 8/3\):

\[\begin{aligned} \frac{\mathrm{d}x}{\mathrm{d}t} &= 10(y-x), \\ \frac{\mathrm{d}y}{\mathrm{d}t} &= x(28-z)-y, \\ \frac{\mathrm{d}z}{\mathrm{d}t} &= xy-\frac{8}{3}z. \end{aligned}\]

After a 100-time-unit spin-up, states are sampled every 0.01 time units. Each of five final splits contains 64 training trajectories of length 20, 32 validation initial conditions, and 128 held-out test initial conditions. Hyperparameters are selected on a separate development split and then frozen. Gaussian noise is added only to training observations, at 0%, 0.1%, 1%, or 5% of each coordinate's clean training standard deviation.

Forecast quality is measured by valid prediction time, or VPT. It ends when normalized squared state error first exceeds 0.4 and is reported in Lyapunov times, \(\lambda_{\max}t\), with \(\lambda_{\max}\approx 0.9134\). A survival curve asks what fraction of forecasts have remained valid up to each time.

Moving the derivative off the data

Strong SINDy estimates time derivatives from observed states and fits a polynomial vector field. That derivative estimate is fragile under noise. Weak SINDy instead fits integrated equations. For a compact test function \(\phi\) and candidate library \(\Theta\), integration by parts gives

\[-\int x(t)\phi'(t)\,\mathrm{d}t =\int\Theta(x(t))\Xi\phi(t)\,\mathrm{d}t.\]

The derivative acts on the known test function rather than the noisy observations. The implementation uses 81-sample windows, a 20-sample stride, four polynomial–Legendre test functions, and trapezoidal quadrature. The weak variants share a degree-two library and the same fitting settings.

Endpoint-weighted weak SINDy leaves each local weak integral unchanged, then weights the resulting regression rows by the window center. The smooth taper is

\[w(s)=\exp\!\left(4-\frac{1}{s(1-s)}\right),\qquad 0<s<1.\]

The endpoint weights are zero and are normalized to unit mean. Both sides of each regression row are multiplied by \(\sqrt{w}\). Uniform weights exactly recover the unweighted weak fit. The frozen threshold is zero, so these benchmark runs test estimation rather than sparse-term pruning; the quadratic library already contains the true Lorenz terms.

The core result

The accepted benchmark gives these restricted mean VPTs over the full valid evaluation horizon. Higher is better.

Restricted mean VPT (LT).
Training noiseStrongWeakWeighted weak
0%8.2686.9116.903
0.1%6.6166.9336.968
1%4.7816.4896.167
5%3.0374.7634.656

At 5% noise, weak SINDy retains about 1.73 more Lyapunov times than strong SINDy. Their 95% bootstrap intervals are 4.623–4.915 and 2.824–3.241 LT. Weighted weak SINDy reaches 4.656 LT, with an interval of 4.383–5.060 LT.

The supported claim is weak-form noise robustness. Endpoint weighting is nearly neutral on clean data, slightly ahead at 0.1% noise, and behind the unweighted weak fit at 1% and 5%. This benchmark does not support a consistent advantage from the extra taper.

What the new curves add

Integral-matching Lorenz receives the exact equation form but not its parameters. It pools cumulative trapezoidal integral equations across training trajectories and solves one through-origin regression for each of \(\sigma\), \(\rho\), and \(\beta\). This avoids derivative estimation, but it has more structural information than state-only SINDy.

Mean restricted VPT (LT) under different information contracts.
MethodInput or noiseVPT
Integral matching0 / 0.1 / 1 / 5%8.684 / 9.046 / 6.950 / 4.120
PandaClean 512-point prefix2.217
Chronos-T5-smallClean 512-point prefix0.549

Integral matching's mean parameter-relative error rises from 0.000089 on clean training data to 0.008950 at 5% noise. The small increase in VPT from 0% to 0.1% lies within estimator and split variation; it is not evidence that noise helps.

Panda is a multivariate forecaster. The Chronos adapter treats the three coordinates independently, predicts in 64-step blocks, and recursively appends those predictions. Their comparison supports Panda over this particular channel-wise Chronos adapter on canonical Lorenz, not a universal ranking of foundation models. Neither forecast curve is a common-information competitor to SINDy: both receive a clean test prefix, external pretraining, and no target-time fitting.

Can a pretrained model emit the equation?

A later experiment asked for something harder than direct forecasting. A 128-sample prefix was mapped through a frozen 1,536-dimensional Panda embedding and an ordered one-layer GRU. A fusion network then emitted support and coefficient heads for a ten-term quadratic library across three state equations: one fixed autonomous SINDy model per prefix, with no target-time optimization or sparse-regression solve.

The architecture was trained on 1,024 trajectories from four non-Lorenz source families and selected on 128 Sprott-B trajectories. The matched ablation changed only the equation loss and a finite-horizon occupation-measure penalty:

Source-validation normalized rollout MSE. Lower is better.
Training objectiveMSEBest epoch
Strong residual1.345825
Weak residual1.1921235
Strong + Birkhoff-MMD1.345825
Weak + Birkhoff-MMD1.19057260

The raw-GRU reference was 1.21226. Most of the improvement came from the derivative-free weak equation; the Birkhoff-MMD term added only 0.13% beyond weak form alone. Here Birkhoff-MMD is a tapered, finite-horizon comparison of predicted and true occupation measures. It is not invariant-measure recovery, and it is not Birkhoff averaging inside Panda. It is also distinct from the endpoint weighting above: that taper weights weak-regression rows, while this one weights forecast samples inside an offline training loss.

This was a one-seed comparison. All four arms eventually encountered non-finite source loss, so selection retained the best earlier finite checkpoint. Every arm also used Panda; the experiment therefore does not isolate a causal contribution from the foundation embedding.

The selected model then received one frozen audit on 16 reserved Lorenz prefixes. Every rollout was finite, but its mean normalized MSE was 2.38643 versus 2.28153 for copying the final context state: a ratio of 1.04598, or 4.60% worse. The source and held-out MSEs use slightly different normalization, so each should be read against its own registered baseline rather than compared directly. The gate failed, so the experiment stopped before CTF4Science. This is target-time zero-update inference, not verified dataset-exclusion zero-shot learning, because Panda's pretraining exposure is unknown.

What I take from it

  • When SINDy is fitted to noisy Lorenz data with the correct library, weak equations provide the clearest robustness gain.
  • The matched endpoint taper is an informative negative result: it does not consistently improve the weak fit.
  • Integral matching is strong when the equation form is known, but that is a different prior-information contract.
  • Foundation models can forecast dynamics directly, yet turning their representations into one stable global equation remains much harder.
  • The amortized weak objective helped across unseen source systems, but honest held-out gating showed that the improvement did not transfer to Lorenz.

Code and evidence: benchmark repository, core Lorenz results, survival extension, and conditioner audit. Local data: core table, extension summary, conditioner summary, and provenance.