By Interestana AI Editorial — AI-drafted, human-overseen. How we report
JEPA-Anything Framework Achieves Domain-Agnostic World Models
Researchers from PhAI Labs, CUHK, Fudan, Stanford, Oxford, and Princeton have introduced JEPA-Anything, a novel domain-agnostic framework designed for constructing world models. This innovative approach departs from traditional methods that require the development of specialized predictive models for each distinct field. Instead, JEPA-Anything employs a single, unified learning recipe that can be applied effectively across a wide array of systems. The framework builds upon existing joint-embedding predictive architectures (JEPAs), such as I-JEPA and V-JEPA 2, by integrating a new technique known as Orthogonal Predictive Factorization (OPF). The research team rigorously tested the capabilities of JEPA-Anything across seven diverse domains: vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather.
The core challenge JEPA-Anything addresses is the "capacity-allocation problem" inherent in standard JEPAs. These standard models typically consist of a context encoder, an EMA target encoder, and a single predictor that generates a monolithic target embedding. This monolithic output can lead to issues where high-variance structures dominate the learning process, causing weaker modes to receive conflicting gradients. Orthogonal Predictive Factorization (OPF) tackles this by decomposing the latent target of width 'd' into 'K' learned subspaces, each of width 'r', such that d = K × r. In most of the experiments conducted, the researchers utilized K = 4. Each of these factors is then assigned its own dedicated predictor, allowing for more specialized learning.
Following the prediction phase, the factor predictions are recombined to form a single, complete latent state. This unified state is then utilized for decoding, planning, or rollout operations. To ensure the utility and stability of these learned factors, the OPF method incorporates three key regularizers. The Orthogonality loss enforces that columns within each projector are orthonormal and that different projectors operate within non-overlapping subspaces. The Factor-activity loss applies a hinge to the per-coordinate standard deviation, preventing any factor from becoming inactive or "dead." Lastly, the Encoder-variance loss provides a direct anti-collapse signal to the online encoder, further stabilizing the learning process. The OPF loss is seamlessly integrated by being added to the original training loss of each respective domain.
Domain-specific adapters are responsible for handling the tokenization and encoder components, ensuring compatibility with diverse data types. The core functionality of the shared learning mechanism is exposed through the OrthogonalFactorProjection library. The importance of orthogonality for stable synthesis was demonstrated in experiments. For instance, on the CITRIS Interventional Pong benchmark, a capacity-matched unconstrained multi-head model exhibited a condition number of 438.52, indicating potential instability, whereas the orthogonal approach aims to mitigate such issues by promoting more structured and stable latent representations across different predictive factors.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.