HeadlinesBriefing favicon HeadlinesBriefing.com

Goldblatt’s Dataset Quantifies Gary Marcus’s AI Skepticism

Hacker News •
×

David Goldblatt released a dataset that catalogs every testable claim Gary Marcus made on Substack from May 2022 to March 2026. The collection contains 2,218 distinct assertions extracted from 474 posts. Goldblatt paired two large‑language‑model pipelines—Claude and ChatGPT—to score each claim against evidence available as of March 2, 2026.

Across the dataset, 59.9% of checkable claims received a 'supported' verdict, 33.7% were 'mixed', and 6.4% were contradicted. Marcus shines on technical details: LLM security flaws hit 100% support, Sora video unreliability 90%, and premature agent deployment 88%, with no contradictions in the AI safety domain.

Marcus falters most in market forecasts. His claim that the GenAI bubble would burst was contradicted 27% of the time, making it his weakest cluster among 54. Predictions shifted from a