KO
Kunle Olukotun
cs.LGcs.ARcs.PLstat.MLcs.DBcs.CLcs.DCcs.AIcs.PFmath.OC
On Valency
published · living versionsW_u7fa5rax·v1 · currentpublished
MLSys: The New Frontier of Machine Learning Systems
with Alexander Ratner, Dan Alistarh, Gustavo Alonso, David G. Andersen +64
1 version
Preprints & journals
51 papers in the corpus · 2011–2026PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX2608.17379v3 · Genghan Zhang, Yixin Dong, Chengze Fan et al.2026 · 0 citationsarXiv
Combee: Scaling Prompt Learning for Self-Improving Language Model Agents2604.04247v2 · Hanchen Li, Runyuan He, Qizheng Zhang et al.2026 · 0 citationsarXiv
Sieve: Dynamic Expert-Aware PIM Acceleration for Evolving Mixture-of-Experts Models2605.11277v1 · Jungwoo Kim, Rubens Lacouture, Genghan Zhang et al.2026 · 0 citationsarXiv
AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization2511.15915v2 · Genghan Zhang, Shaowei Zhu, Anjiang Wei et al.2025 · 0 citationsarXiv
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models2510.04618v3 · Qizheng Zhang, Changran Hu, Shubhangi Upasani et al.2025 · 3 citationsarXiv
AI+HW 2035: Shaping the Next Decade2603.05225v1 · Deming Chen, Jason Cong, Azalia Mirhoseini et al.2026 · 0 citationsarXiv
Streaming Tensor Programs: A Streaming Abstraction for Dynamic Parallelism2511.07776v2 · Gina Sohn, Genghan Zhang, Konstantin Hossfeld et al.2025 · 1 citationarXiv
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents2506.14852v2 · Qizheng Zhang, Michael Wornow, Gerry Wan et al.2025 · 0 citationsarXiv
FuseFlow: A Fusion-Centric Compilation Framework for Sparse Deep Learning on Streaming Dataflow2511.04768v2 · Rubens Lacouture, Nathan Zhang, Ritvik Sharma et al.2025 · 1 citationarXiv
Cyclotron: Compilation of Recurrences to Distributed and Systolic Architectures2511.09987v1 · Shiv Sundram, Akhilesh Balasingam, Nathan Zhang et al.2025 · 0 citationsarXiv
Adaptive Self-improvement LLM Agentic System for ML Library Development2502.02534v2 · Genghan Zhang, Weixin Liang, Olivia Hsu et al.2025 · 1 citationarXiv
SSM-RDU: A Reconfigurable Dataflow Unit for Long-Sequence State-Space Models2503.22937v2 · Sho Ko, Kunle Olukotun2025 · 0 citationsarXiv
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits2502.08141v1 · Zikai Zhou, Qizheng Zhang, Hermann Kumbong et al.2025 · 0 citationsarXiv
DFModel: Design Space Optimization of Large-Scale Systems Exploiting Dataflow Mappings2412.16432v1 · Sho Ko, Nathan Zhang, Olivia Hsu et al.2024 · 0 citationsarXiv
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts2405.07518v2 · Raghu Prabhakar, Ram Sivaramakrishnan, Darshan Gandhi et al.2024 · 34 citationsarXiv
Implementing and Optimizing the Scaled Dot-Product Attention on Streaming Dataflow2404.16629v2 · Gina Sohn, Nathan Zhang, Kunle Olukotun2024 · 4 citationsarXiv
Revet: A Language and Compiler for Dataflow Threads2302.06124v2 · Alexander Rucker, Shiv Sundram, Coleman Smith et al.2023 · 4 citationsarXiv
BaCO: A Fast and Portable Bayesian Compiler Optimization Framework2212.11142v2 · Erik Hellsten, Artur Souza, Johannes Lenfers et al.2022 · 22 citationsarXiv
The Sparse Abstract Machine2208.14610v2 · Olivia Hsu, Maxwell Strange, Ritvik Sharma et al.2022 · 39 citationsASPLOS 2023: Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems Volume 3 (2023) 710-726
Stardust: Compiling Sparse Tensor Algebra to a Reconfigurable Dataflow Architecture2211.03251v1 · Olivia Hsu, Alexander Rucker, Tian Zhao et al.2022 · 9 citationsarXiv
Homunculus: Auto-Generating Efficient Data-Plane ML Pipelines for Datacenter Networks2206.05592v1 · Tushar Swamy, Annus Zulfiqar, Luigi Nardi et al.2022 · 26 citationsarXiv
Efficient Memory Partitioning in Software Defined Hardware2202.01261v3 · Matthew Feldman, Tian Zhao, Kunle Olukotun2022 · 1 citationarXiv
Taurus: A Data Plane Architecture for Per-Packet ML2002.08987v2 · Tushar Swamy, Alexander Rucker, Muhammad Shahbaz et al.2020 · 94 citationsarXiv
Capstan: A Vector RDA for Sparsity2104.12760v2 · Alexander Rucker, Matthew Vilim, Tian Zhao et al.2021 · 37 citationsarXiv
Bayesian Optimization with a Prior for the Optimum2006.14608v4 · Artur Souza, Luigi Nardi, Leonardo B. Oliveira et al.2020 · 27 citationsarXiv
EmptyHeaded: A Relational Engine for Graph Processing.28077912 · Aberger, Christopher R, Tu, Susan, Olukotun, Kunle et al.2020 · 13 citationsProceedings. ACM-SIGMOD International Conference on Management of Data. 2016;2016:431-446
Ensuring Rapid Mixing and Low Bias for Asynchronous Gibbs Sampling.28344730 · De Sa, Christopher, Olukotun, Kunle, Ré, Christopher2020 · 28 citationsJMLR workshop and conference proceedings. 2016;48:1567-1576
Exploring the Utility of Developer Exhaust.31131381 · Zhang, Jian, Lam, Max, Wang, Stephanie et al.2020 · 1 citationProceedings of the Second Workshop on Data Management for End-to-End Machine Learning. Workshop on Data Management for End-to-End Machine Learning (2nd : 2018 : Houston, Tex.). 2018;2018
Understanding and Optimizing Asynchronous Low-Precision Stochastic Gradient Descent.29391770 · De Sa, Christopher, Feldman, Matthew, Ré, Christopher et al.2020 · 106 citationsProceedings. International Symposium on Computer Architecture. 2017;2017:561-574
Analysis of DAWNBench, a Time-to-Accuracy Machine Learning Performance Benchmark1806.01427v2 · Cody Coleman, Daniel Kang, Deepak Narayanan et al.2018 · 107 citationsarXiv
MLSys: The New Frontier of Machine Learning Systems1904.03257v3 · Alexander Ratner, Dan Alistarh, Gustavo Alonso et al.2019 · 20 citationsarXivon Valency
Rapidly Mixing Gibbs Sampling for a Class of Factor Graphs Using Hierarchy Width.27279724 · De Sa, Christopher, Zhang, Ce, Olukotun, Kunle et al.2019 · 15 citationsAdvances in neural information processing systems. 2015;28:3079-3087
Taming the Wild: A Unified Analysis of Hogwild!-Style Algorithms.27330264 · De Sa, Christopher, Zhang, Ce, Olukotun, Kunle et al.2019 · 114 citationsAdvances in neural information processing systems. 2015;28:2656-2664
Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator1909.13654v1 · Tian Zhao, Yaqi Zhang, Kunle Olukotun2019 · 10 citationsProceedings of the 2 nd SysML Conference, Palo Alto, CA, USA, 2019. Copyright 2019 by the author(s)
Practical Design Space Exploration1810.05236v3 · Luigi Nardi, David Koeplinger, Kunle Olukotun2018 · 76 citationsarXiv
Efficient Multiway Hash Join on Reconfigurable Hardware1905.13376v1 · Kunle Olukotun, Raghu Prabhakar, Rekha Singhal et al.2019 · 0 citationsarXiv
Polystore++: Accelerated Polystore System for Heterogeneous Workloads1905.10336v1 · Rekha Singhal, Nathan Zhang, Luigi Nardi et al.2019 · 5 citationsICDCS 2019
DeepFreak: Learning Crystallography Diffraction Patterns with Automated Machine Learning1904.11834v2 · Artur Souza, Leonardo B. Oliveira, Sabine Hollatz et al.2019 · 7 citationsarXiv
High-Accuracy Low-Precision Training1803.03383v1 · Christopher De Sa, Megan Leszczynski, Jian Zhang et al.2018 · 69 citationsarXiv
LevelHeaded: Making Worst-Case Optimal Joins Work in the Common Case1708.07859v1 · Christopher R. Aberger, Andrew Lamb, Kunle Olukotun et al.2017 · 0 citationsarXiv
Infrastructure for Usable Machine Learning: The Stanford DAWN Project1705.07538v2 · Peter Bailis, Kunle Olukotun, Christopher Re et al.2017 · 16 citationsarXiv
Flare: Native Compilation for Heterogeneous Workloads in Apache Spark1703.08219v1 · Gr'egory M. Essertel, Ruby Y. Tahboub, James M. Decker et al.2017 · 10 citationsarXiv
EmptyHeaded: A Relational Engine for Graph Processing1503.02368v7 · Christopher R. Aberger, Susan Tu, Kunle Olukotun et al.2015 · 13 citationsarXiv
Ensuring Rapid Mixing and Low Bias for Asynchronous Gibbs Sampling1602.07415v3 · Christopher De Sa, Kunle Olukotun, Christopher R'e2016 · 34 citationsarXiv
Old Techniques for New Join Algorithms: A Case Study in RDF Processing1602.03557v1 · Christopher R. Aberger, Susan Tu, Kunle Olukotun et al.2016 · 26 citationsarXiv
Generating Configurable Hardware from Parallel Patterns1511.06968v1 · Raghu Prabhakar, David Koeplinger, Kevin Brown et al.2015 · 57 citationsarXiv
Taming the Wild: A Unified Analysis of Hogwild!-Style Algorithms1506.06438v2 · Christopher De Sa, Ce Zhang, Kunle Olukotun et al.2015 · 114 citationsarXiv
Rapidly Mixing Gibbs Sampling for a Class of Factor Graphs Using Hierarchy Width1510.00756v1 · Christopher De Sa, Ce Zhang, Kunle Olukotun et al.2015 · 16 citationsarXiv
Global Convergence of Stochastic Gradient Descent for Some Non-convex Matrix Problems1411.1134v3 · Christopher De Sa, Kunle Olukotun, Christopher R'e2014 · 102 citationsarXiv
Utilizing Static Analysis and Code Generation to Accelerate Neural Networks1206.6466v1 · Lawrence McAfee, Kunle Olukotun2012 · 1 citationarXiv
Career total: 314 works. 51 are in this corpus.Showing the 50 most recent.
Profile built from the corpus for this byline.
Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.