Goodfire推出探针式智能体监控器,成本较传统方案大幅降低
专注于AI可解释性的初创公司Goodfire于本周推出一款面向AI智能体的新型监控方案,通过读取模型内部运行信号来识别异常行为,而非仅依赖外部模型对输出内容进行复核。
该监控器目前面向模型托管服务商Baseten的客户开放,可监测攻击性黑客行为、生化武器滥用及奖励黑客等风险。客户可根据需要选择日志记录、人工审核或直接拒绝请求等自动响应方式。
Goodfire的机制类似机场安检:小型探测器(probes)在智能体工作过程中不断读取模型内部信号,只有在探测到异常时,才会调用独立AI模型进行详细检查。由于探针复用了模型前向传播过程中已有的计算,相比逐一读取智能体全部输出的传统监控方式,运行成本显著降低。
Goodfire在Kimi K3上的测试显示,监控约1500次会话的探针调用成本约为51美元,而采用较便宜的AI模型逐步检查需要约233美元,使用顶级模型则需约1万美元。测试中探针捕获了94%的恶意黑客会话,并将8.7%的无害会话送入二次检查。同时运行四个探针给模型响应启动时间带来的增加不到2%。
- 2026-10-09 00:00 | TechCrunch:Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost阅读原文
The standard way to keep an AI agent in line is to have a second AI read over its shoulder. It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text. Goodfire, a startup focused on interpretability (figuring ...