SB

Serge Belongie

cs.CVcs.LGcs.AIcs.CLstat.MLcs.CRcs.GRAlgorithmscs.NEeess.IV

On Valency

published · living versions
W_eyra28jp·v1 · currentpublished
Adaptively Learning the Crowd Kernel
with Omer Tamuz, Ce Liu, Ohad Shamir, Adam Tauman Kalai
1 version

Preprints & journals

154 papers in the corpus · 2004–2026
Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models2607.26326v2 · Jiaang Li, Chengzu Li, Zhaochong An et al.2026 · 0 citationsarXiv
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities2607.25948v1 · Mingqiao Ye, Zhaochong An, Zhitong Gao et al.2026 · 0 citationsarXiv
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward2603.26599v2 · Zhaochong An, Orest Kupyn, Th'eo Uscidda et al.2026 · 1 citationarXiv
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training2602.06285v3 · Lucia Gordon, Serge Belongie, Christian Igel et al.2026 · 0 citationsarXiv
RAIGen: Rare Attribute Identification in Text-to-Image Generative Models2602.06806v3 · Silpa Vadakkeeveetil Sreelatha, Dan Wang, Serge Belongie et al.2026 · 0 citationsarXiv
Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks2606.15534v1 · Feng Qiao, Zhaochong An, Zhexiao Xiong et al.2026 · 0 citationsarXiv
The Latent Color Subspace: Emergent Order in High-Dimensional Chaos2603.12261v2 · Mateusz Pach, Jessica Bader, Quentin Bouniot et al.2026 · 0 citationsarXiv
Stitched Value Model for Diffusion Alignment2605.19804v1 · Hyojun Go, Hyungjin Chung, Prune Truong et al.2026 · 0 citationsarXiv
Beyond Binary Success: A Diagnostic Meta-Evaluation Framework for Fine-Grained Manipulation2605.19986v1 · He-Yang Xu, Pengyuan Zhang, Zongyuan Ge et al.2026 · 0 citationsarXiv
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution2512.09115v2 · Sander Riis\oen Jyhne, Christian Igel, Morten Goodwin et al.2025 · 0 citationsarXiv
Video Understanding: From Geometry and Semantics to Unified Models2603.17840v1 · Zhaochong An, Zirui Li, Mingqiao Ye et al.2026 · 3 citationsMachine Intelligence Research 2026
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding2505.14462v2 · Jiaang Li, Yifei Yuan, Wenyan Li et al.2025 · 0 citationsarXiv
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning2601.21037v1 · Chengzu Li, Zanyi Wang, Jiaang Li et al.2026 · 0 citationsarXiv
Better Language Models Exhibit Higher Visual Alignment2410.07173v3 · Jona Ruthardt, Gertjan J. Burghouts, Serge Belongie et al.2024 · 0 citationsarXiv
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models2504.02821v3 · Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot et al.2025 · 4 citationsarXiv
RespoDiff: Dual-Module Bottleneck Transformation for Responsible & Faithful T2I Generation2509.15257v2 · Silpa Vadakkeeveetil Sreelatha, Sauradip Nag, Muhammad Awais et al.2025 · 0 citationsarXiv
What if Othello-Playing Language Models Could See?2507.14520v2 · Xinyi Chen, Yifei Yuan, Jiaang Li et al.2025 · 0 citationsarXiv
Revealing Fine-Grained Values and Opinions in Large Language Models2406.19238v3 · Dustin Wright, Arnav Arora, Nadav Borenstein et al.2024 · 0 citationsarXiv
Noise-Coded Illumination for Forensic and Photometric Video Analysis2507.23002v1 · Peter F. Michael, Zekun Hao, Serge Belongie et al.2025 · 4 citationsACM Trans. Graph. 44, 5, Article 165 (October 2025), 16 pages
Multi-Modal Framing Analysis of News2503.20960v3 · Arnav Arora, Srishti Yadav, Maria Antoniak et al.2025 · 4 citationsarXiv
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model2503.16282v2 · Zhaochong An, Guolei Sun, Yun Liu et al.2025 · 12 citationsarXiv
POEM: Precise Object-level Editing via MLLM control2504.08111v1 · Marco Schouten, Mehmet Onurcan Kaya, Serge Belongie et al.2025 · 0 citationsarXiv
Taxonomy-Aware Evaluation of Vision-Language Models2504.05457v1 · V'esteinn Sn\aebjarnarson, Kevin Du, Niklas Stoehr et al.2025 · 0 citationsarXiv
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis2502.18180v2 · Lei Li, Sen Jia, Jianhao Wang et al.2025 · 0 citationsarXiv
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation2410.22489v4 · Zhaochong An, Guolei Sun, Yun Liu et al.2024 · 0 citationsarXiv
Unlearning-based Neural Interpretations2410.08069v2 · Ching Lam Choi, Alexandre Duplessis, Serge Belongie2024 · 0 citationsChoi, Ching Lam, Alexandre Duplessis, and Serge Belongie. 'Unlearning-Based Neural Interpretations'. In The Thirteenth International Conference on Learning Representations, 2025
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes2501.13851v1 · Shiling Deng, Serge Belongie, Peter Ebert Christensen2025 · 0 citationsarXiv
Discriminative Class Tokens for Text-to-Image Diffusion Models2303.17155v4 · Idan Schwartz, V'esteinn Sn\aebjarnarson, Hila Chefer et al.2023 · 10 citationsarXiv
Familiarity-Based Open-Set Recognition Under Adversarial Attacks2311.05006v2 · Philip Enevoldsen, Christian Gundersen, Nico Lang et al.2023 · 0 citationsarXiv
LoQT: Low-Rank Adapters for Quantized Pretraining2405.16528v4 · Sebastian Loeschcke, Mads Toftrup, Michael J. Kastoryano et al.2024 · 0 citationsarXiv
MMEarth: Exploring Multi-Modal Pretext Tasks For Geospatial Representation Learning2405.02771v2 · Vishal Nedungadi, Ankit Kariryaa, Stefan Oehmcke et al.2024 · 42 citationsarXiv
Labeled Data Selection for Category Discovery2406.04898v2 · Bingchen Zhao, Nico Lang, Serge Belongie et al.2024 · 2 citationsarXiv
Coarse-To-Fine Tensor Trains for Compact Visual Representations2406.04332v1 · Sebastian Loeschcke, Dan Wang, Christian Leth-Espensen et al.2024 · 0 citationsarXiv
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction2212.10564v3 · Boyi Li, Rodolfo Corona, Karttikeya Mangalam et al.2022 · 2 citationsarXiv
An algorithm competition for automatic species identification from herbarium specimens.32626608 · Little, Damon P, Tulig, Melissa, Tan, Kiat Chuan et al.2024 · 42 citationsApplications in plant sciences. 2020;8(6):e11365
The Plant Pathology Challenge 2020 data set to classify foliar disease of apples.33014634 · Thapa, Ranjita, Zhang, Kai, Snavely, Noah et al.2024 · 304 citationsApplications in plant sciences. 2020;8(9):e11390
Rethinking Few-shot 3D Point Cloud Semantic Segmentation2403.00592v1 · Zhaochong An, Guolei Sun, Yun Liu et al.2024 · 48 citationsarXiv
Learning to Taste: A Multimodal Wine Dataset2308.16900v4 · Thoranna Bender, Simon Moe S\orensen, Alireza Kashani et al.2023 · 6 citationsarXiv
Assessing Neural Network Robustness via Adversarial Pivotal Tuning2211.09782v2 · Peter Ebert Christensen, V'esteinn Sn\aebjarnarson, Andrea Dittadi et al.2022 · 0 citationsarXiv
Prompt, Condition, and Generate: Classification of Unsupported Claims with In-Context Learning2309.10359v1 · Peter Ebert Christensen, Srishti Yadav, Serge Belongie2023 · 0 citationsarXiv
The Herbarium 2021 Half-Earth Challenge Dataset and Machine Learning Competition.35178056 · de Lutio, Riccardo, Park, John Y, Watson, Kimberly A et al.2023 · 35 citationsFrontiers in plant science. 2021;12:787127
Three New Validators and a Large-Scale Benchmark Ranking for Unsupervised Domain Adaptation2208.07360v4 · Kevin Musgrave, Serge Belongie, Ser-Nam Lim2022 · 1 citationarXiv
Fashionpedia-Ads: Do Your Favorite Advertisements Reveal Your Fashion Taste?2305.02360v1 · Mengyun Shi, Claire Cardie, Serge Belongie2023 · 0 citationsarXiv
Fashionpedia-Taste: A Dataset towards Explaining Human Fashion Taste2305.02307v1 · Mengyun Shi, Serge Belongie, Claire Cardie2023 · 0 citationsarXiv
Polynomial Neural Fields for Subband Decomposition and Manipulation2302.04862v1 · Guandao Yang, Sagie Benaim, Varun Jampani et al.2023 · 8 citationsarXiv
SITTA: Single Image Texture Translation for Data Augmentation2106.13804v2 · Boyi Li, Yin Cui, Tsung-Yi Lin et al.2021 · 3 citationsarXiv
Career total: 405 works. 154 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.