This blog post is based on recent work "Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks," published at COLM 2026 and written with Yunxiang Zhang and Professor Lu Wang at the University of Michigan. The post addresses the disconnect between AI investments in scientific discovery and agent performance on real ML research challenges, noting breakthroughs like Alpha Evolve and claims of solving Navier-Stokes. The authors argue creativity offers a useful lens, defined by Mark A.
Runco and Garrett J. Jaeger as the production of ideas that are simultaneously original and useful. Drawing from creative psychology, the definition breaks creativity into four subcomponents: P-Creativity (novelty relative to program history), H-Creativity (novelty compared to human knowledge), Impact, and Feasibility.
The main question analyzes whether performance differences between agent frameworks can be attributed to how they structure and guide the creative search process. The post further breaks originality down into P-Creativity and H-Creativity following Boden, and usefulness into impact and feasibility following Chan and Schunn. The framework aims to quantify how creativity emerges and evolves within different agent scaffolding systems to explain performance gaps in ML engineering tasks.
Source: Towards Data Science · Résumé par HeadlinesBriefing