The new system functions like airport security, using lightweight detectors to scan a model’s internal signals as it processes data. Only when a probe identifies a potential threat—such as unauthorized hacking or weapon-related queries—does a secondary model intervene for a deeper inspection. By tapping into calculations the model is already performing, Goodfire reduces the computational overhead that plagues standard, external monitoring tools.
In performance testing using the Kimi K3 model, the system processed 1,500 sessions for approximately $51, significantly undercutting the $233 to $10,000 range required by conventional monitoring methods. The probes achieved a 94% detection rate for malicious hacking attempts while adding less than 2% to the model's response time. Customers using the Baseten hosting platform can now configure these monitors to trigger specific responses, ranging from simple event logging to total request termination.
This technology addresses a growing security gap in the open-source AI ecosystem, where models are frequently stripped of safety guardrails. As these agents gain the ability to interact with external environments, Goodfire CTO Dan Balsam argues that inference-time safeguards are becoming a necessity rather than an option. The company plans to build on this foundation by reverse-engineering how specific behaviors emerge during model training, moving toward a future of precision-engineered AI safety.

Comments (0)
No comments yet. Be the first!