Skip to content
@FoundationVision

FoundationVision

Bytedance's opensource FoundationVision models

Welcome to FoundationVision @ ByteDance!

Introduction šŸ‘‹

Hello! This is the GitHub space for the FoundationVision @ ByteDance.

We are dedicated to exploring the frontiers of multimodal intelligence, with the ultimate goal of building Artificial General Intelligence systems (AGI).

Our research focuses on deep learning and multimodal intelligence. We are particularly interested in:

  • Visual Foundation Models, Generative Pretrained Models and Large Language Models.
  • Multimodal Foundation Models and Representation Learning.
  • Open World Interaction via Unified Multi-modal generation and understanding.
  • Large-scale Multi-modal generative Pretraining and Alignment.

Our group strives to push the boundaries of multimodal intelligence and has produced highly influential works in the field, including:

Popular repositories Loading

  1. VAR VAR Public

    [NeurIPS 2024 Best Paper Award][GPT beats diffusionšŸ”„] [scaling laws in visual generationšŸ“ˆ] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". A…

    Jupyter Notebook 8.7k 573

  2. ByteTrack ByteTrack Public

    [ECCV 2022] ByteTrack: Multi-Object Tracking by Associating Every Detection Box

    Python 6.7k 1.1k

  3. LlamaGen LlamaGen Public

    Autoregressive Model Beats Diffusion: šŸ¦™ Llama for Scalable Image Generation

    Python 2k 95

  4. Infinity Infinity Public

    [CVPR 2025 Oral]Infinity āˆž : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

    Python 1.6k 93

  5. GLEE GLEE Public

    [CVPR2024 Highlight]GLEE: General Object Foundation Model for Images and Videos at Scale

    Python 1.2k 77

  6. Waver Waver Public

    Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.

    953 124

Repositories

Showing 10 of 21 repositories
  • Liquid Public

    (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators

    FoundationVision/Liquid's past year of commit activity
    Python 640 MIT 35 12 0 Updated Jun 1, 2026
  • InfinityStar Public

    [NeurIPS 2025 Oral]Infinityā­ļø: Unified Spacetime AutoRegressive Modeling for Visual Generation

    FoundationVision/InfinityStar's past year of commit activity
    Python 785 MIT 28 8 0 Updated Apr 16, 2026
  • Infinity Public

    [CVPR 2025 Oral]Infinity āˆž : Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

    FoundationVision/Infinity's past year of commit activity
    Python 1,587 MIT 93 61 5 Updated Apr 16, 2026
  • Alive Public

    [Tech Report] Alive: A Unified Audio-Video Generation Model

    FoundationVision/Alive's past year of commit activity
    457 30 7 0 Updated Mar 31, 2026
  • .github Public
    FoundationVision/.github's past year of commit activity
    0 0 0 0 Updated Nov 20, 2025
  • UniTok Public

    [NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding

    FoundationVision/UniTok's past year of commit activity
    Python 529 MIT 14 16 1 Updated Nov 14, 2025
  • VAR Public

    [NeurIPS 2024 Best Paper Award][GPT beats diffusionšŸ”„] [scaling laws in visual generationšŸ“ˆ] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!

    FoundationVision/VAR's past year of commit activity
    Jupyter Notebook 8,730 MIT 573 57 (1 issue needs help) 5 Updated Nov 10, 2025
  • Waver Public

    Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.

    FoundationVision/Waver's past year of commit activity
    953 124 16 2 Updated Aug 27, 2025
  • BitVAE Public

    official training and inference code of bitwise tokenizer

    FoundationVision/BitVAE's past year of commit activity
    Python 72 MIT 2 4 0 Updated May 18, 2025
  • GenerateU Public

    [CVPR2024] Generative Region-Language Pretraining for Open-Ended Object Detection

    FoundationVision/GenerateU's past year of commit activity
    Python 196 MIT 9 15 0 Updated Mar 29, 2025

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…