One predictive core, seven different worlds
The team behind JEPA-Anything showed that one predictive architecture can learn the dynamics of cells, molecules, fluids and planets without being told the laws of physics.

Machine learning usually splits into specialities. One network predicts the next video frames. Another predicts the movements of a robotic arm. Another predicts the weather. Yet another predicts the trajectories of molecules. The team behind the paper "JEPA-Anything" asked the opposite question: is there one shared principle for building a predictive model that works across radically different systems?
What a world model is
In this framing, a world model is any network that builds an internal state from the information available and uses that state to predict another state of the same system. That "other state" can be the future in time, a hidden part of space, a different point of view, or the effect of an intervention. The JEPA family (Joint-Embedding Predictive Architecture) predicts in representation space rather than in raw pixels. The model's capacity then goes into structure that actually predicts something, not into reproducing textures and noise.
The problem shows up when different scales and rates of change land in one shared prediction target. The easier, stronger signals then take up most of the capacity, and subtle but important changes get pushed out. The authors answer this with a mechanism called OPF (orthogonal predictive factorization). It cuts the target state into complementary subspaces, each handled by a separate prediction branch, and then merges them back. A bundle of constraints, namely orthogonality of the factors, their activity and the variance of the encoder, is meant to keep the branches from learning the same thing and to protect against representation collapse. The factors are not assigned in advance to specific physical quantities. What they represent is decided by the structure of the data.
A test on seven systems
The authors did not stop at video or control. They checked the same predictive core in seven domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather. Each domain keeps its own way of observing, its own construction of the context-target pair and its own encoder. What is shared is the prediction layer and a single hidden-state interface.
In a controlled environment, an intervention version of Pong, the model was trained only on data with single changes and tested on unseen combinations of factors. On the training distribution the prediction error fell by about 11.7 percent relative to standard JEPA. On unseen combinations the mean squared error dropped by about 3.5 percent, and the improvement appeared in all five random draw pairs.
In the "Matched Dynamics Benchmark" both architectures received identical data, encoder, state-transition backbone, compute budget and evaluation split. JEPA-Anything improved the result in nine of ten tasks, among others by 39.7 percent on the Burgers equation, by 39.3 percent on the shallow water equation and by 10.5 percent on WeatherBench 2. Long step-by-step predictions, in which the model's own results go back into the input, were tested separately.
The most telling example concerns astronomy. The model was given simulated positions and velocities of planets without being told Kepler's laws. An analysis of internal prediction patterns revealed a dependence of orbital frequency on semi-major axis with a slope of -1.4991 against the theoretical value of -1.5 for Kepler's third law, with a fit coefficient R² of 0.99999999.
In biology, the factors learned by the model were used to propose an intervention combining cytokines with antibody blockade. The idea went through further stages of verification: on cell lines, on patient-derived organoids and in mice. Molecular dynamics was checked on liquid water, alpha quartz, paracetamol and benzene. The authors' conclusion is cautious but clear: it is not about replacing domain knowledge, but about a shared layer of prediction learning on which that knowledge can be built.
Sources
2- 01量子位: AI离“理解万物”还有多远?先拿癌细胞和行星轨道试试水ZH
- 02JEPA-Anything: Learning Predictive Models across Different WorldsEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.