cs.LGcs.CLcs.AIcs.SIstat.MLcs.CVcs.IRcs.CYphysics.soc-phAlgorithms

On Valency

published · living versions
W_3seyyh2p·v1 · currentpublished
Evaluating Large Language Models Trained on Code
with Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan +53
1 version

Preprints & journals

144 papers in the corpus · 2006–2026
SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators2608.07641v1 · Yuheng Zhang, Yuanchun Wang, Fanjin Zhang et al.2026 · 0 citationsarXiv
GSPNet: Graph Spectral Projection Network Using Learnable Spectral Transformation.42013247 · Geng, Yangli-Ao, Dong, Yuxiao, Feng, Wenzheng et al.2026 · 0 citationsIEEE transactions on pattern analysis and machine intelligence. 2026;48(9):10248-10265
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability2512.05394v3 · Shizhan Liu, Xinran Deng, Zhuoyi Yang et al.2025 · 0 citationsarXiv
Glyph: Scaling Context Windows via Visual-Text Compression2510.17800v2 · Jiale Cheng, Yusen Liu, Xinyu Zhang et al.2025 · 0 citationsarXiv
AgentBench: Evaluating LLMs as Agents2308.03688v3 · Xiao Liu, Hao Yu, Hanchen Zhang et al.2023 · 53 citationsarXiv
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning2509.02544v2 · Haoming Wang, Haoyang Zou, Huatong Song et al.2025 · 0 citationsarXiv
GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model2412.03930v2 · Yunhe Pang, Bo Chen, Fanjin Zhang et al.2024 · 1 citationarXiv
ProtGO: universal protein function prediction utilizing multi-modal gene ontology knowledge.40632605 · Wang, Boyan, Geng, Yangliao, Cheng, Xingyi et al.2025 · 11 citationsBioinformatics (Oxford, England). 2025;41(7)
UavNetSim-v1: A Python-based Simulation Platform for UAV Communication Networks2507.09852v1 · Zihao Zhou, Zipeng Dai, Linyi Huang et al.2025 · 8 citationsarXiv
Colony Binary Classification Based on Persistent Homology Feature Extraction and Improved EfficientNet.40564441 · Wang, Zumin, Yang, Ke, Tang, Jie et al.2025 · 0 citationsBioengineering (Basel, Switzerland). 2025;12(6)
Deep learning-based automated segmentation for the quantitative diagnosis of cerebral small vessel disease via multisequence MRI.40496122 · Zhao, Huiyu, Zhang, Miaoyi, Tang, Weijun et al.2025 · 6 citationsFrontiers in neurology. 2025;16:1540923
HPSS: Heuristic Prompting Strategy Search for LLM Evaluators2502.13031v2 · Bosi Wen, Pei Ke, Yufei Sun et al.2025 · 3 citationsarXiv
xTrimoPGLM: unified 100-billion-parameter pretrained transformer for deciphering the language of proteins.40181110 · Chen, Bo, Cheng, Xingyi, Li, Pan et al.2025 · 56 citationsNature methods. 2025;22(5):1028-1039
CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer2408.06072v3 · Zhuoyi Yang, Jiayan Teng, Wendi Zheng et al.2024 · 18 citationsarXiv
StepMathAgent: A Step-Wise Agent for Evaluating Mathematical Processes through Tree-of-Error2503.10105v1 · Shu-Xun Yang, Cunxiang Wang, Yidong Wang et al.2025 · 0 citationsarXiv
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning2402.04236v3 · Ji Qi, Ming Ding, Weihan Wang et al.2024 · 2 citationsarXiv
Small Language Model Makes an Effective Long Text Extractor2502.07286v1 · Yelin Chen, Fanjin Zhang, Jie Tang2025 · 2 citationsarXiv
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks2412.15204v2 · Yushi Bai, Shangqing Tu, Jiajie Zhang et al.2024 · 24 citationsarXiv
Dynamic Scaling of Unit Tests for Code Reward Modeling2501.01054v1 · Zeyao Ma, Xiaokang Zhang, Jing Zhang et al.2025 · 1 citationarXiv
CogAgent: A Visual Language Model for GUI Agents2312.08914v3 · Wenyi Hong, Weihan Wang, Qingsong Lv et al.2023 · 179 citationsarXiv
The Superalignment of Superhuman Intelligence with Large Language Models2412.11145v2 · Minlie Huang, Yingkang Wang, Shiyao Cui et al.2024 · 1 citationarXiv
MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model2409.13729v2 · Zhen Yang, Jinhao Chen, Zhengxiao Du et al.2024 · 0 citationsarXiv
ReST-MCTS*: LLM Self-Training via Process Reward Guided Tree Search2406.03816v3 · Dan Zhang, Sining Zhoubian, Ziniu Hu et al.2024 · 27 citationsarXiv
AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents2410.24024v2 · Yifan Xu, Xiao Liu, Xueqiao Sun et al.2024 · 3 citationsarXiv
AutoWebGLM: A Large Language Model-based Web Navigating Agent2404.03648v2 · Hanyu Lai, Xiao Liu, Iat Long Iong et al.2024 · 44 citationsarXiv
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models2408.15778v4 · Jiayi Gui, Yiming Liu, Jiale Cheng et al.2024 · 1 citationarXiv
CogVLM2: Visual Language Models for Image and Video Understanding2408.16500v1 · Wenyi Hong, Weihan Wang, Ming Ding et al.2024 · 7 citationsarXiv
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs2408.07055v1 · Yushi Bai, Jiajie Zhang, Xin Lv et al.2024 · 2 citationsarXiv
Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed Graphs2407.15431v1 · Huanjing Zhao, Beining Yang, Yukuo Cen et al.2024 · 14 citationsarXiv
Does Negative Sampling Matter? a Review With Insights Into its Theory and Applications.38421844 · Yang, Zhen, Ding, Ming, Huang, Tinglin et al.2024 · 54 citationsIEEE transactions on pattern analysis and machine intelligence. 2024;46(8):5692-5711
OAG-Bench: A Human-Curated Benchmark for Academic Graph Mining2402.15810v2 · Fanjin Zhang, Shijie Shi, Yifan Zhu et al.2024 · 17 citationsProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD '24), August 25--29, 2024, Barcelona, Spain
LGB: Language Model and Graph Neural Network-Driven Social Bot Detection2406.08762v2 · Ming Zhou, Dan Zhang, Yuandong Wang et al.2024 · 17 citationsarXiv
Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer2405.04312v2 · Zhuoyi Yang, Heyang Jiang, Wenyi Hong et al.2024 · 4 citationsarXiv
Algorithm xxx: Faster Randomized SVD with Dynamic Shifts2404.09276v1 · Xu Feng, Wenjian Yu, Yuyang Xie et al.2024 · 0 citationsarXiv
BOND: Bootstrapping From-Scratch Name Disambiguation with Multi-task Promoting2404.08322v1 · Yuqing Cheng, Bo Chen, Fanjin Zhang et al.2024 · 5 citationsProceedings of TheWebConf 2024 (WWW '24), May 13--17, 2024, Singapore
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion2403.05121v1 · Wendi Zheng, Jiayan Teng, Zhuoyi Yang et al.2024 · 23 citationsarXiv
Does Negative Sampling Matter? A Review with Insights into its Theory and Applications2402.17238v1 · Zhen Yang, Ming Ding, Tinglin Huang et al.2024 · 54 citationsarXiv
Tuning Multiple Counter-Anions in Porous Coordination Polymers with lcy Topology for Acetylene/Ethylene Separation.38335451 · Tang, Jie, Shen, Yuebing, He, Xingge et al.2024 · 6 citationsInorganic chemistry. 2024;63(8):3667-3674
RecDCL: Dual Contrastive Learning for Recommendation2401.15635v2 · Dan Zhang, Yangliao Geng, Wenwen Gong et al.2024 · 63 citationsProceedings of TheWebConf 2024 (WWW '24), May 13--17, 2024, Singapore
SketchNE: Embedding Billion-Scale Networks Accurately in One Hour2110.12782v3 · Yuyang Xie, Yuxiao Dong, Jiezhong Qiu et al.2021 · 14 citationsarXiv
Career total: 626 works. 144 are in this corpus.Showing the 50 most recent.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.