HeadlinesBriefing favicon HeadlinesBriefing.com

Consumer Inference Systems: Mobile AI Benchmarks

Hacker News •
×

Artificial Analysis publishes Consumer Inference Systems benchmarks evaluating AI model performance on mobile hardware, specifically the iPhone 17 Pro. The study uses a simple average of five evaluations — BFCL (subset), IFBench, AA-Omniscience, GPQA Diamond, and MATH-500 — chosen to represent real-world mobile device usage. Context is limited to 16K tokens maximum.

Intelligence scores measure model capability across tool calling, instruction following, knowledge accuracy, non-hallucination rate, scientific reasoning, and quantitative reasoning. Higher scores indicate better performance.

End-to-end generation time tracks total wall-clock seconds to process a 1,024-token prompt and generate a 256-token response. Lower times are better. Results reflect fundamental hardware and architecture constraints, excluding verbosity effects.

Token efficiency tracks context budget overruns — generations stopping at the 16K-token limit instead of completing. Models too large for the device or exceeding time limits are omitted from performance measurements.