Architecture Name
Mobius
Parent issue
#1
Motivations
We propose Mobius, an architecture that provides stronger support than Transformer for several prominent contemporary directions, including self-evolving intelligence, world models, scientific discovery, and hardware–software co-design.
Proposed Architecture
Mobius achieves these goals by disentangling knowledge and reasoning within the base model. Consequently, Mobius natively supports novel capabilities unavailable to Transformer, such as flexible residual connections, latent reasoning, and continuous latent-space diffusion......
Preliminary Results (if any)
Our preliminary results indicate that Mobius-v0.0 achieves approximately 1.6× faster convergence and nearly 4× end-to-end inference speedup compared to Transformer. Beyond these fundamental capability gains, validation of Mobius on additional high-value scenarios remains ongoing.
Experiments Plan
We plan to announce, within the next few days, our 30B-parameter model trained via continual pre-training (CPT), along with the accompanying technical report. This CPT model is built on the Mobius-v0.0 architecture. Over the coming months, we will progressively release higher versions of our architecture (v0.5, v1.0, and beyond), together with the training-from-scratch results on these variants.
Architecture Name
Mobius
Parent issue
#1
Motivations
We propose Mobius, an architecture that provides stronger support than Transformer for several prominent contemporary directions, including self-evolving intelligence, world models, scientific discovery, and hardware–software co-design.
Proposed Architecture
Mobius achieves these goals by disentangling knowledge and reasoning within the base model. Consequently, Mobius natively supports novel capabilities unavailable to Transformer, such as flexible residual connections, latent reasoning, and continuous latent-space diffusion......
Preliminary Results (if any)
Our preliminary results indicate that Mobius-v0.0 achieves approximately 1.6× faster convergence and nearly 4× end-to-end inference speedup compared to Transformer. Beyond these fundamental capability gains, validation of Mobius on additional high-value scenarios remains ongoing.
Experiments Plan
We plan to announce, within the next few days, our 30B-parameter model trained via continual pre-training (CPT), along with the accompanying technical report. This CPT model is built on the Mobius-v0.0 architecture. Over the coming months, we will progressively release higher versions of our architecture (v0.5, v1.0, and beyond), together with the training-from-scratch results on these variants.