GS

Giulio Starace

cs.CLcs.AIcs.CYcs.LGcs.CVcs.SDeess.AS

On Valency

published · living versions
W_6cysh3yv·v1 · currentpublished
PaperBench: Evaluating AI's Ability to Replicate AI Research
with Oliver Jaffe, Dane Sherburn, James Aung, Jun Shern Chan +8
1 version

Preprints & journals

4 papers in the corpus · 2022–2024
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering2410.07095v6 · Jun Shern Chan, Neil Chowdhury, Oliver Jaffe et al.2024 · 9 citationsarXiv
GPT-4o System Card2410.21276v1 · OpenAI: Aaron Hurst, Adam Lerer, Adam P. Goucher et al.2024 · 165 citationsarXiv
Probing LLMs for Joint Encoding of Linguistic Categories2310.18696v1 · Giulio Starace, Konstantinos Papakostas, Rochelle Choenni et al.2023 · 3 citationsarXiv
[Re] Badder Seeds: Reproducing the Evaluation of Lexical Methods for Bias Measurement2206.01767v1 · Jille van der Togt, Lea Tiyavorabun, Matteo Rosati et al.2022 · 0 citationsRescience C, 8(2), #40; 2022
Career total: 5 works. 4 are in this corpus.

Profile built from the corpus for this byline.

Author records are still filling in while the Hub is in alpha. If this is your page, you'll be able to claim it soon. Spot a mistake? Tell us.