DP

Dimitris Papailiopoulos

cs.LGstat.MLcs.DCcs.ITmath.ITcs.AIcs.CLmath.OCcs.DScs.NI

On Valency

published · living versions
W_u7fa5rax·v1 · currentpublished
MLSys: The New Frontier of Machine Learning Systems
with Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G. Andersen +64
1 version

Preprints & journals

80 papers in the corpus · 2010–2026
Scaling Discovery through Test-Time Communication2609.21032v1 · Jongho Park, Vasilis Kontonis, Shivam Garg et al.2026 · 0 citationsarXiv
You Don't Need to Run Every Eval2606.24020v2 · Yuchen Zeng, Dimitris Papailiopoulos2026 · 0 citationsarXiv
Wait, Wait, Wait... Why Do Reasoning Models Loop?2512.12895v2 · Charilaos Pipis, Shivam Garg, Vasilis Kontonis et al.2025 · 0 citationsarXiv
SuperThoughts: Reasoning Tokens in Superposition2606.13862v1 · Zheyang Xiong, Shivam Garg, Max Yu et al.2026 · 0 citationsarXiv
Provable Deterministic Leverage Score Sampling1404.1530v3 · Dimitris Papailiopoulos, Anastasios Kyrillidis, Christos Boutsidis2014 · 68 citationsarXiv
Sparse Principal Component of a Rank-deficient Matrix1106.1651v1 · Megasthenis Asteris, Dimitris S. Papailiopoulos, George N. Karystinos2011 · 14 citationsarXiv
ECHO: Terminal Agents Learn World Models for Free2605.24517v1 · Vaishnavi Shrivastava, Piero Kauffmann, Ahmed Awadallah et al.2026 · 0 citationsarXiv
Open-World Evaluations for Measuring Frontier AI Capabilities2605.20520v1 · Sayash Kapoor, Peter Kirgis, Andrew Schwartz et al.2026 · 0 citationsarXiv
MEMENTO: Teaching LLMs to Manage Their Own Context2604.09852v1 · Vasilis Kontonis, Yuchen Zeng, Shivam Garg et al.2026 · 0 citationsarXiv
Endless Terminals: Scaling RL Environments for Terminal Agents2601.16443v3 · Kanishk Gandhi, Shivam Garg, Noah D. Goodman et al.2026 · 0 citationsarXiv
ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning2512.00831v2 · Yuchen Zeng, Shuibai Zhang, Wonjun Kang et al.2025 · 0 citationsarXiv
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models2510.10964v1 · Junhyuck Kim, Ethan Ewer, Taehong Moon et al.2025 · 0 citationsarXiv
Extrapolation by Association: Length Generalization Transfer in Transformers2506.09251v2 · Ziyang Cai, Nayoung Lee, Avi Schwarzschild et al.2025 · 0 citationsarXiv
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data2502.06737v2 · Thomas Zeng, Shuibai Zhang, Shutong Wu et al.2025 · 0 citationsarXiv
Phi-4-reasoning Technical Report2504.21318v1 · Marah Abdin, Sahaj Agarwal, Ahmed Awadallah et al.2025 · 1 citationarXiv
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges2502.01612v2 · Nayoung Lee, Ziyang Cai, Avi Schwarzschild et al.2025 · 1 citationarXiv
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit2501.09240v1 · Liu Yang, Ziqian Lin, Kangwook Lee et al.2025 · 0 citationsarXiv
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries2412.08890v1 · Junhyuck Kim, Jongho Park, Jaewoong Cho et al.2024 · 1 citationarXiv
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition2410.05603v1 · Zheyang Xiong, Ziyang Cai, John Cooper et al.2024 · 0 citationsarXiv
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding2307.05908v2 · Seongjun Yang, Gibbeum Lee, Jaewoong Cho et al.2023 · 1 citationarXiv
CHAI: Clustered Head Attention for Efficient LLM Inference2403.08058v2 · Saurabh Agarwal, Bilge Acun, Basil Hosmer et al.2024 · 2 citationsarXiv
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks2402.04248v2 · Jongho Park, Jaeseung Park, Zheyang Xiong et al.2024 · 10 citationsarXiv
Looped Transformers are Better at Learning Learning Algorithms2311.12424v3 · Liu Yang, Kangwook Lee, Robert Nowak et al.2023 · 0 citationsarXiv
How Well Can Transformers Emulate In-context Newton's Method?2403.03183v1 · Angeliki Giannou, Liu Yang, Tianhao Wang et al.2024 · 0 citationsarXiv
Dissecting Chain-of-Thought: Compositionality through In-Context Filtering and Learning2305.18869v2 · Yingcong Li, Kartik Sreenivasan, Angeliki Giannou et al.2023 · 2 citationsarXiv
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation2305.04533v1 · Gibbeum Lee, Volker Hartmann, Jongho Park et al.2023 · 33 citationsarXiv
Mini-Batch Optimization of Contrastive Loss2307.05906v1 · Jaewoong Cho, Kartik Sreenivasan, Keon Lee et al.2023 · 2 citationsarXiv
Teaching Arithmetic to Small Transformers2307.03381v1 · Nayoung Lee, Kartik Sreenivasan, Jason D. Lee et al.2023 · 7 citationsarXiv
PathProx: A Proximal Gradient Algorithm for Weight Decay Regularized Deep Neural Networks2210.03069v4 · Liu Yang, Jifan Zhang, Joseph Shenouda et al.2022 · 0 citationsarXiv
The Expressive Power of Tuning Only the Normalization Layers2302.07937v2 · Angeliki Giannou, Shashank Rajput, Dimitris Papailiopoulos2023 · 0 citationsarXiv
Cuttlefish: Low-Rank Model Training without All the Tuning2305.02538v2 · Hongyi Wang, Saurabh Agarwal, Pongsakorn U-chupala et al.2023 · 0 citationsarXiv
Transformers as Algorithms: Generalization and Stability in In-context Learning2301.07067v2 · Yingcong Li, M. Emrullah Ildiz, Dimitris Papailiopoulos et al.2023 · 11 citationsarXiv
Looped Transformers as Programmable Computers2301.13196v1 · Angeliki Giannou, Shashank Rajput, Jy-yong Sohn et al.2023 · 3 citationsarXiv
Utilizing Language-Image Pretraining for Efficient and Robust Bilingual Word Alignment2205.11616v2 · Tuan Dinh, Jy-yong Sohn, Shashank Rajput et al.2022 · 0 citationsarXiv
LIFT: Language-Interfaced Fine-Tuning for Non-Language Machine Learning Tasks2206.06565v4 · Tuan Dinh, Yuchen Zeng, Ruisu Zhang et al.2022 · 55 citationsarXiv
Rare Gems: Finding Lottery Tickets at Initialization2202.12002v2 · Kartik Sreenivasan, Jy-yong Sohn, Liu Yang et al.2022 · 12 citationsarXiv
GenLabel: Mixup Relabeling using Generative Models2201.02354v1 · Jy-yong Sohn, Liang Shang, Hongxu Chen et al.2022 · 2 citationsarXiv
Permutation-Based SGD: Is Random Optimal?2102.09718v2 · Shashank Rajput, Kangwook Lee, Dimitris Papailiopoulos2021 · 1 citationarXiv
Finding Everything within Random Binary Networks2110.08996v2 · Kartik Sreenivasan, Shashank Rajput, Jy-yong Sohn et al.2021 · 1 citationarXiv
On the Utility of Gradient Compression in Distributed Training Systems2103.00543v3 · Saurabh Agarwal, Hongyi Wang, Shivaram Venkataraman et al.2021 · 18 citationsarXiv
An Exponential Improvement on the Memorization Capacity of Deep Threshold Networks2106.07724v1 · Shashank Rajput, Kartik Sreenivasan, Dimitris Papailiopoulos et al.2021 · 9 citationsarXiv
Optimal Lottery Tickets via SubsetSum: Logarithmic Over-Parameterization is Sufficient2006.07990v2 · Ankit Pensia, Shashank Rajput, Alliot Nagle et al.2020 · 35 citationsarXiv
Pufferfish: Communication-efficient Models At No Extra Cost2103.03936v1 · Hongyi Wang, Saurabh Agarwal, Dimitris Papailiopoulos2021 · 5 citationsarXiv
Bad Global Minima Exist and SGD Can Reach Them1906.02613v2 · Shengchao Liu, Dimitris Papailiopoulos, Dimitris Achlioptas2019 · 37 citationsarXiv
Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification2010.16248v1 · Saurabh Agarwal, Hongyi Wang, Kangwook Lee et al.2020 · 13 citationsarXiv
Attack of the Tails: Yes, You Really Can Backdoor Federated Learning2007.05084v1 · Hongyi Wang, Kartik Sreenivasan, Shashank Rajput et al.2020 · 109 citationsarXiv
Closing the convergence gap of SGD without replacement2002.10400v6 · Shashank Rajput, Anant Gupta, Dimitris Papailiopoulos2020 · 9 citationsarXiv
DETOX: A Redundancy-based Framework for Faster and More Robust Gradient Aggregation1907.12205v2 · Shashank Rajput, Hongyi Wang, Zachary Charles et al.2019 · 47 citationsarXiv
Federated Learning with Matched Averaging2002.06440v1 · Hongyi Wang, Mikhail Yurochkin, Yuekai Sun et al.2020 · 97 citationsarXiv
Career total: 125 works. 80 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.