HeadlinesBriefing HeadlinesBriefing 2 languages

Where Does the Money Go Across Long-Running Coding Agents?

Towards Data Science ·

🇬🇧 English

Straight Up AI, a small consultancy, has built an internal control plane for orchestrating coding agents. How much latitude those agents get depends on three factors: the blast radius of a mistake, project context, and project maturity. The team argues that one factor is missing: the cost of building. Running the control plane through a Claude Max account costs £200 a month, so the marginal cost of bad decisions is mainly time. Before developing it further, the team needed to check whether it could genuinely be afforded.

To find out, the team analysed seven weeks of activity, from 15 July to 4 September, covering 44 development cycles across its portfolio. Converting every token to input-token equivalents, with cached reads at 0.1x, cached writes at 1.25x and output tokens at 5x, showed that API rates would cost 22 times more than the current utilisation. As the consultancy grows, that gap quickly becomes untenable and approaches the cost of a mid-level engineer's salary.

Time was the second concern. Adversarial review, which spins up a second agent to challenge every commit, was the suspected cause of slow runs. The data largely supported this, though not for the reason first assumed. Reviewers were dispatched at a ratio of 1.19x the number of implementers, and 26% of them requested changes. Each change triggered a remediation agent and another review cycle, which risked rabbit-holing rather than prompting a step back to consider structural changes beyond the commit's scope.

The control plane uses one controller that lives for the whole job and dispatches one worker per increment. The controller acts as an independent arbiter, validating whether each increment is completed, reviewed or failed, which prevents incorrect progress claims from cascading through the plan.

View original article →


🇨🇳 简体中文

长期运行的编码智能体,钱都花在哪里?

Straight Up AI 是一家小型咨询公司,它为编排编码智能体搭建了一个内部控制平面。智能体可获得多大的自主权,取决于三个因素:错误的影响范围、项目背景和项目成熟度。该团队认为,还缺少一个因素:构建成本。通过 Claude Max 账户运行该控制平面每月需花费 200 英镑,因此糟糕决策的边际成本主要体现为时间。在进一步开发之前,团队需要确认它是否真的负担得起。

为此,团队分析了从 7 月 15 日到 9 月 4 日共七周的运行情况,涵盖其项目组合中的 44 个开发周期。将每个 token 换算为输入 token 当量,其中缓存读取按 0.1 倍、缓存写入按 1.25 倍、输出 token 按 5 倍计算,结果显示,按 API 费率计费的成本将是现有使用量的 22 倍。随着咨询公司的壮大,这一差距很快会变得难以承受,其金额接近一名中级工程师的薪资。

时间是第二个问题。对抗性审查会启动第二个智能体来挑战每一次提交,它被怀疑是运行缓慢的原因。数据在很大程度上支持了这一判断,不过原因与最初的设想并不相同。审查者的派遣数量是实施者的 1.19 倍,其中 26% 的审查者要求修改。每次修改都会触发一个修复智能体和另一轮审查,这可能导致陷入细枝末节,而不是退一步考虑超出提交范围的结构性改动。

控制平面采用一个贯穿整个任务的控制器,并为每个增量分派一个工作智能体。控制器充当独立的仲裁者,验证每个增量是已完成、已审查还是已失败,从而防止错误的进度声明在整个计划中层层蔓延。

简体中文 version →