Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Scenic: A Jax Library for Computer Vision Research and Beyond
| Date | Stars |
|---|---|
| 2026-07-24 | 3819 |
| 2026-07-25 | 3819 |
| 2026-07-28 | 3818 |
| 2026-07-30 | 3818 |
| 2026-07-31 | 3818 |
| 2026-08-06 | 3818 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Scenic <div style="text-align: left"> <img align="right" src="https://raw.githubusercontent.com/google-research/scenic/main/images/scenic_logo.png" width="200" alt="Scenic logo"></img> </div> *Scenic* is a codebase with a focus on research around attention-based models for computer vision. Scenic has been successfully used to develop classification, segmentation, and detection models for multiple modalities including images, video, audio, and multimodal combinations of them. More precisely, *Scenic* is a (i) set of shared light-weight libraries solving tasks commonly encountered tasks when training large-scale (i.e. multi-device, multi-host) vision models; and (ii) several *projects* containing fully fleshed out problem-specific training and evaluation loops using these libraries. Scenic is developed in [JAX](https://github.com/jax-ml/jax) and uses [Flax](https://github.com/google/flax). ### Contents * [What we offer](#what-we-offer) * [SOTA models and baselines in Scenic](#sota-models-and-baselines-in-scenic) * [Philosophy](#philosophy) * [Getting started](#getting-started) * [Scenic component design](#scenic-component-design) * [Citing Scenic](#citing-scenic) ## What we offer Among others *Scenic* provides * Boilerplate code for launching experiments, summary writing, logging, profiling, etc; * Optimized training and evaluation loops, losses, metrics, bi-partite matchers, etc; * Input-pipelines for popular vision datasets; * [Baseline models](https://github.com/google-research/scenic/tree/main/scenic/projects/baselines#scenic-baseline-models), including strong non-attentional baselines. ## SOTA models and baselines in *Scenic* There are some SOTA models and baselines in Scenic which were either developed using Scenic, or have been reimplemented in Scenic: Projects that were developed in Scenic or used it for their experiments: * [ViViT: A Video Vision Transformer](https://arxiv.org/abs/2103.15691) * [OmniNet: Omnidirectional Representations from Transformers](https://arxiv.org/abs/2103.01075) * [Attention Bottlenecks for Multimodal Fusion](https://arxiv.org/abs/2107.00135) * [TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?](https://arxiv.org/abs/2106.11297) * [Exploring the Limits of Large Scale Pre-training](https://arxiv.org/abs/2110.02095) * [The Efficiency Misnomer](https://arxiv.org/abs/2110.12894) * [Discrete Representations Strengthen Vision Transformer Robustness](https://arxiv.org/abs/2111.10493) * [Pyramid Adversarial Training Improves ViT Performance](https://arxiv.org/abs/2111.15121) * [VUT: Versatile UI Transformer for Multi-Modal Multi-Task User Interface Modeling](https://arxiv.org/abs/2112.05692) * [CLAY: Learning to Denoise Raw Mobile UI Layouts for Improving Datasets at Scale](https://arxiv.org/abs/2201.04100) * [Zero-Shot Text-Guided Object Generation with Dream Fields](https://arxiv.org/abs/2112.01455) * [Multiview Transformers for Video Recognition](https://arxiv.org/abs/2201.04288) * [PolyViT: Co-training Vision Transformers on Images, Videos and Audio](https://arxiv.org/abs/2111.12993) * [Simple Open-Vocabulary Object Detection with Vision Transformers](https://arxiv.org/abs/2205.06230) * [Learning with Neighbor Consistency for Noisy Labels](https://arxiv.org/abs/2202.02200) * [Token Turing Machines](https://arxiv.org/pdf/2211.09119.pdf) * [Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning](https://arxiv.org/pdf/2302.14115.pdf) * [AVATAR: Unconstrained Audiovisual Speech Recognition](https://arxiv.org/abs/2206.07684) * [Adaptive Computation with Elastic Input Sequence](https://arxiv.org/abs/2301.13195) * [Location-Aware Self-Supervised Transformers for Semantic Segmentation](https://arxiv.org/abs/2212.02400) * [How can objects help action recognition?](https://openaccess.thecvf.com/content/CVPR2023/html/Zhou_How_Can_Objects_Help_Action_Recognition_CVPR_2023_paper.html) * [Verbs in Action: Improving verb understanding in video-langu
Excerpt of 12,142 characters
Read on GitHub168
66
59
Peter Hawkins · Google
33
19
Alexey Gritsenko · Netherlands
19
Hana Joo · https://www.linkedin.com/in/hana-joo-6a0379127/ · Germany
17
Thomas Unterthiner · Google
16
Utku Evci · @google · Canada
10
manoj kumar · Google Brain · Netherlands
9
Marcus Chiam · Google DeepMind · United Kingdom
8
Santiago Castro · @Netflix · United States
7
6
Yilei · United States
6
Rebecca Chen
6
5
4
4
4
4
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:c7647caadd7a3fa1, topic:deep-learning, topic:jax, readme:pretraining
matched fp:c7647caadd7a3fa1, topic:computer-vision, desc:computer vision, readme:computer vision