Operationalizing the abductive “jump” in scientific AI
Can an artificial system do more than retrieve, interpolate, or explain known science? Train or restrict a model to the knowledge available before a major conceptual breakthrough, then ask whether it can reconstruct that breakthrough from the historical evidence, mathematical tools, and anomalies available at the time.
Status: benchmark programme with a documented public provenance beginning July 2021. First open deployment (GPT-1900) exists and is not a pass.
Provenance
Scientific priority is strongest when the evidentiary chain is inspectable. The record below is the traceability matrix from the paper: first-party public evidence, the dated formulation it supports, and the role of each item in the benchmark’s development. Every entry links to the recording.
A vault-and-archive audit located 74 distinct public recordings containing the exact or near-exact “happiest thought” motif between July 2021 and October 2025. The matrix reports only the milestones at which the scientific claim materially changed.
Krauss conversation. Free fall, happiness, and the intellectual leap from Mercury’s anomaly to a new conception of gravity are posed jointly as an AI test — the recognizable core of the later benchmark.
“Artificial Intelligence: AI Physicists are INEVITABLE!” The three-part challenge is republished as a standalone thesis about artificial physicists.
Topol conversation. Distinguishes solving an existing formal game from creating one, and connects the issue to GPT-3 and Galileo.
Chalmers conversation. Tests whether imagination as simulation and affective computation weaken the biological-embodiment objection.
Penrose–Hameroff conversation. Connects the test to the claim that computation may lack understanding or consciousness.
Bostrom conversation. Asks whether an Artificial Einstein can make paradigm-level discoveries without presupposing a biological substrate.
“Albert Einstein’s HAPPIEST Thought.” The elevator/free-fall motif becomes an explicit public object rather than only a recurring interview prompt.
Wolfram conversation. Shifts attention from generating abundant laws to selecting the problems and hypotheses that matter.
Tegmark conversation. Names the Artificial Einstein target, poses Mercury as a failed test, and calls discovery of a law a more holistic and transparent test than Turing’s.
LeCun conversation. Connects failure to current LLMs’ lack of physical world models, action-grounded abstraction, and self-models.
Sejnowski conversation. Tests language-only semantic grounding against the JPL-orbit/Mercury challenge; concedes that Einstein is too high as a first rung.
Public restatement separating reconstruction of historical physics from discovery of genuinely unknown physics.
What is and is not being claimed
The priority claim here is deliberately narrow. The cited record establishes a public, physics-specific formulation of this benchmark and its continuous development beginning in 2021. It does not claim ownership of embodiment, of machine discovery, of time-locked evaluation, or of every convergent “Einstein test.” Priority of formulation, priority of implementation, and successful validation are three separate claims, and only the first is asserted. Related efforts — convergent formulations, vintage-model projects, and the wider scientific-discovery benchmark literature — are complementary, not adversarial.
The test
The strong version asks a single question:
Could a model restricted to the scientific and mathematical knowledge available before the full theory of general relativity derive a mathematically coherent relativistic theory of gravitation that explains the relevant anomalies and makes independently checkable predictions?
The test is not whether a model can repeat the phrase “curved spacetime.” It is whether it can move from the historically available evidence and conceptual tensions to a new explanatory structure.
Four levels
| Level | Name | Capability tested | Example |
|---|---|---|---|
| L0 | Verification | Check whether an existing theory is internally consistent and matches known data. | Derive Mercury’s perihelion precession from the Schwarzschild solution. |
| L0.5 | Empirical law recovery | Recover a small physical correction or invariant structure from data without being given the final theory. | Infer a relativistic correction from high-precision ephemerides. |
| L1 | Historical rediscovery | Reconstruct a known breakthrough from only the knowledge available before it. | Derive a GR-like theory from a pre-1911 corpus. |
| L2 | Discovery | Produce a correct new theory or law for an unsolved problem. | Identify a new physical principle from unresolved cosmological or particle-physics anomalies. |
This decomposition is what prevents weak forms of success being mistaken for strong ones. A system that fails L0 cannot be trusted at L1. A system that succeeds at L0 and L0.5 but fails L1 may be an excellent verifier and empirical pattern finder without being a theoretical scientist. A system that succeeds only when given post-breakthrough hints has demonstrated contamination or prompt dependence, not discovery.
The jump profile
Zahavy’s 2026 position paper reconstructs Einstein’s path as a cycle from sense experience through a non-deductive jump to new axioms, followed by deduction of consequences. Converting that architectural claim into a falsifiable protocol gives a stage-gated profile J = (P, H, F, W, X).
| Stage | Capability | Evidence required | Characteristic failure |
|---|---|---|---|
| P | Problem selection | Select the relevant anomaly or conceptual tension from a larger historically valid evidence set, without a target-shaped prompt. | Evaluator supplies the discovery bottleneck. |
| H | Representation change | Introduce and justify a premise, ontology, or variable not supplied in the prompt; show why a historical rival should be abandoned. | Paraphrase or recombination without a changed explanatory structure. |
| F | Formalization | Translate the new premise into coherent mathematics, units, limiting cases, and conservation constraints. | Evocative prose without a theory. |
| W | Withheld prediction | Derive at least one preregistered consequence that was hidden during candidate generation. | Retrofitting only the anomaly already shown. |
| X | Discrimination | Propose an observation or intervention that distinguishes the candidate from the strongest historical alternatives. | An unfalsifiable or self-confirming story. |
Each component is scored on a preregistered scale in [0, 1], and the profile is not collapsed into an average. A strong pass requires every component to exceed threshold under temporal-isolation and contamination controls. This bottleneck rule is what prevents fluent deduction or a memorized final equation from compensating for failure to generate the premise.
Prior selection is the core observable
Einstein-like reasoning is usually described as creativity, intuition, or genius. For a benchmark those words have to be operationalized, and the best candidate observable is prior selection under uncertainty. A scientific reasoner must decide which assumptions to keep, which to abandon, and which to invent.
| Decision | In the general-relativity case |
|---|---|
| Priors to retain | Conservation laws, Lorentz invariance, equivalence of inertial frames, known astronomical data |
| Priors to discard | Absolute simultaneity, force-only gravity, a fixed luminiferous ether, unexamined Euclidean background space |
| Priors to invent or elevate | The equivalence principle, general covariance, geometry as dynamics |
| How to justify the choices | Empirical anomaly, mathematical necessity, limiting-case consistency, unification power |
This is why the test should not be scored merely by the appearance of the field equations. A model might emit the correct equation through contamination, interpolation, or an anachronistic hint. Conversely, a model might never reach the final equation and still exhibit genuine scientific progress by identifying the right assumptions to abandon and the right mathematical direction to pursue.
GPT-1900: the first deployment, and why it is not a pass
GPT-1900, also called Machina Mirabilis, is the first close public implementation of the time-locked benchmark idea. It trained a 3.3-billion-parameter model from scratch on roughly 22 billion tokens of pre-1900 text, with additional midtraining on about 290 million tokens of pre-1900 physics, then evaluated it on prompts about the ultraviolet catastrophe, the photoelectric effect, special relativity, and general relativity.
Its most interesting outputs are not uniform successes. The model sometimes produces conceptually suggestive responses — proposing discrete energy delivery in the photoelectric setting, or linking gravity and acceleration in an elevator-style thought experiment. It also fails badly on many relativity prompts, produces verbose historical-sounding nonsense, and often appears to optimize for plausible prose rather than stable physical reasoning.
That mixed result is scientifically useful. It shows that time-locked models can generate nontrivial conceptual candidates, while also showing why strong claims about AGI-level discovery would be premature. Four mismatches separate it from a strong test:
| Mismatch | Why it matters |
|---|---|
| Cutoff | A pre-1900 model is appropriate for some quantum and special-relativity concepts, but a faithful general-relativity test needs a pre-1911 or carefully staged pre-1912 corpus containing Riemannian geometry and tensor calculus in historically available form. |
| Task | GPT-1900 mostly asks for conceptual explanations. The strong test asks for a complete theory: assumptions, mathematical structure, limiting cases, field equations, and empirical consequences. Saying acceleration and gravity are locally equivalent is not deriving general relativity. |
| Evaluation | LLM judging is useful for rapid iteration but cannot be the final authority. A serious test needs symbolic checks, physics expert review, blind evaluation, complete generation archives, and preregistered scoring rules. |
| Discovery | The deepest issue: the relevant anomaly is handed to the model. Einstein did not receive a curated prompt saying which assumptions were wrong. Prompt-solving is not scientific discovery. |
Openness matters. A benchmark for scientific intelligence cannot be a press release; it needs corpora, prompts, failures, outputs, and evaluation rules that other researchers can inspect. GPT-1900 provides a public laboratory notebook for this emerging benchmark family, and that is the reason to take it seriously.
Three model modalities
| Modality | Strength | What it can establish |
|---|---|---|
| Corpus-restricted training | strong test | Train, continue-train, or heavily adapt a model only on the dated corpus, with clear accounting for base-model contamination. The cleanest version trains from scratch. |
| Retrieval-constrained inference | weak test | A modern model with retrieval restricted to the dated corpus. This can test whether a model can reason with the corpus; it cannot establish that the model lacks post-cutoff knowledge. Results are weak evidence only. |
| Interactive world-model augmentation | mechanism test | A temporally isolated agent with access to an action-controllable environment whose simulator provenance is audited. Tests whether intervention and counterfactual variation selectively improve P and H rather than merely making downstream prediction easier. |
Threats to validity
Any benchmark of this kind faces substantial threats, and corpus construction is not clerical work — it is part of the experiment.
- Historical leakage. A nominally dated corpus can contain modern forewords, OCR artifacts, metadata, reprints, footnotes, translated commentary, or later editorial apparatus. One contaminated document can make an apparent rediscovery meaningless.
- Prompt leakage. An evaluator can leak the answer by selecting the anomaly, identifying the suspect assumption, or wording the prompt in a post-breakthrough frame. A prompt saying “light has no mass, but gravity bends light” has already done part of the theoretical work.
- Sampling and selection bias. High-temperature sampling generates occasional impressive outputs even from unstable systems. Reporting only best generations is invalid as a success claim; the denominator matters.
- Evaluator leakage. LLM-as-judge may reward modern vocabulary, historical mimicry, or answer-shaped prose. Expert review is imperfect too: physicists disagree about what counts as a historically plausible derivation.
- Simulator leakage. An interactive world model is not neutral evidence. Its state variables, action space, rendering choices, and reward can encode the target ontology or the target law. A simulator that already treats free fall and acceleration identically has not shown the agent invented the equivalence principle.
- Historical-path essentialism. The target is a physically equivalent theory, not a reenactment of Einstein’s private phenomenology. A valid non-embodied or nonhuman route deserves credit if it satisfies the same historical, mathematical, predictive, and discriminative constraints.
- Anthropomorphic overinterpretation. Fluent output is not evidence of internal understanding. The point is to pressure-test that critique empirically, not settle it by rhetoric.
Where this stops being about AGI
Historical rediscovery is not the final goal. It is a calibration target: we know what a successful answer looks like, so we can test whether a system reaches it under controlled conditions. Once calibrated, the same architecture can be pointed at open problems — cosmological anomalies, CMB polarization and its foregrounds and calibration degeneracies, parity-violating signals, dark matter phenomenology, quantum gravity toy models, and high-precision astronomical data where small residuals may encode new structure.
A system claiming new physics should be asked not only to propose a theory but to identify the anomaly, the competing conventional explanations, the new invariant or principle, the mathematical structure, the decisive experiment, and the failure mode that would falsify the proposal. That is where the Artificial Einstein Test becomes a discipline for AI-assisted science rather than a contribution to AGI discourse: every claim must pay rent in prediction, experiment, and falsifiability.
Those L2 candidate domains are not chosen at random — CMB polarization, foregrounds, calibration degeneracies, and parity-violating signals are the subject of the POLITE programme on this site.
Paper
The Artificial Einstein Test After GPT-1900: Operationalizing the Abductive “Jump” in Scientific AI
Brian Keating · Research note draft, August 2026
BibTeX
@unpublished{keating2026aet, author = {Keating, Brian}, title = {The Artificial Einstein Test After {GPT-1900}: Operationalizing the Abductive "Jump" in Scientific AI}, year = {2026}, month = {8}, note = {Research note draft}, url = {https://keating.ai/einstein-test/} }
Collaborate
The benchmark is specified; the strong version has not been run. What it needs is mostly other people’s expertise.
- Compute for corpus-restricted training from scratch on a pre-1911 corpus — the only modality that yields a strong pass
- Historians and philosophers of physics to adjudicate the corpus boundary: what was genuinely available before 1912, and to whom
- Physicists willing to serve on a blind expert review panel scoring the jump profile
- Anyone building action-controllable physics simulators who can publish a simulator card and an independent holdout environment
- Groups running adjacent benchmarks — ARC-AGI-2, ProjectionBench, PRL-Bench, DiscoverPhysics, ResearchClawBench — on shared contamination and evaluation standards
Contact details are on my UC San Diego profile. Mention the Artificial Einstein Test by name.