This blog post summarizes the research 'Can LLM Agents Discover? Evaluating Creativity on ML Engineering Tasks' presented at COLM 2026 by Yunxiang Zhang and Professor Lu Wang from the University of Michigan. It explores why LLM agents, despite claims of breakthroughs like Alpha Evolve and Kosmos AI scientist, still underperform top human researchers in real ML tasks. The authors argue that creativity—defined as the production of ideas that are both original and useful—offers a useful lens to evaluate and improve agent performance.
Drawing from creative psychology, they break creativity into P-Creativity (novelty relative to the agent's memory), H-Creativity (novelty versus all human knowledge), impact, and feasibility. The post examines how agent frameworks structure and guide the creative search process, aiming to quantify how creativity emerges and evolves within different designs. By linking creativity to search in conceptual spaces—referencing Boden, Newell, Shaw, and Simon—the work seeks to understand performance differences between frameworks through their ability to support novel and effective idea generation.
The goal is to develop better evaluation metrics for LLM agents in scientific discovery by measuring creativity as a combination of originality and usefulness.
Source: Towards Data Science · Summarized by HeadlinesBriefing