Core Background & Industry Pain Points Nearly all modern Multimodal Large Language Models (MLLMs)—ranging from frontier proprietary models to mainstream open-source architectures—rely on decoupled pipelines pairing a pretrained visual encoder (e.g., CLIP, SigLIP) with a projection layer and a language model backbone. While this paradigm successfully bootstrap visual capabilities by borrowing frozen visual priors, it creates severe architectural friction: 1. Information Bottlenecks and Cross-Modal Misalignment: Fixed-dimension representations from independently pretrained vision towers cause irreversible semantic loss on high-resolution spatial relationships, dense OCR text, and long-horizon video context; 2. Engineering Pipeline & Distributed Training Overhead: Maintaining separate network topologies with distinct tensor and pipeline parallelisms complicates distributed clusters and introduces serving latency bubbles. While an "encoder-free" architecture ingesting raw pixel patches directly into an autoregressive LLM offers simplicity, its scaling behavior has never been systematically charted. ### Architectural Highlights & Underlying Mechanics Researchers from CASIA and collaborating institutions present the first principled, multi-order scaling laws comparing encoder-free and encoder-based MLLMs across identical compute budgets: 1. Unified Autoregressive Pixel Tokenization: Raw pixel patches are mapped into input embeddings via a lightweight linear projection, entering the Transformer directly as standard tokens without external vision towers; 2. Shifted Compute-Optimal Allocation: While the text loss-compute frontier is virtually indistinguishable between the two paradigms, removing the visual encoder systematically shifts optimal compute allocation toward larger parameter counts for multimodal objectives, whereas text objective allocation remains invariant; 3. Emergent Vision Adaptation in Transformer Layers: As compute scales, the language backbone naturally assumes visual abstraction: early Transformer layers develop bidirectional attention across visual tokens, and sparse MoE routers dynamically direct visual patches to specialized early-stage experts. ### Benchmark & Experimental Validation - 10^22 FLOPs Crossover Threshold: At lower compute scales, encoder-based models lead due to inherited CLIP priors. However, the performance curve of encoder-free models scales at a steeper slope, crossing and surpassing encoder-based baselines at approximately 10^22 FLOPs—well within enterprise pretraining thresholds; - Diminishing Returns of Frozen Priors: The quantitative analysis confirms that the marginal value of pretrained visual encoders degrades exponentially with training compute; - Interleaved Image-Text Coherence: On complex multimodal reasoning benchmarks, the encoder-free architecture maintains consistent attention geometry across mixed text and image sequences, avoiding context fragmentation common to external projection bridges. ### Engineering Takeaways & Practical Guide - Paper & Reference: Full formal formulations and Pareto loss curves are documented in arXiv:2609.35457; - Architectural Recommendations: Large-scale labs training foundation models with pretraining compute exceeding 10^22 FLOPs should actively deprecate decoupled vision towers in favor of unified end-to-end architectures to simplify inference pipelines and training parallelism; - Edge Deployment Considerations: For smaller budget regimes (< 10^21 FLOPs), encoder-based designs remain cost-effective priors, or can serve as distillation targets for unified architectures.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.