Goodfire claims its new "inside-out" monitoring tool can capture out-of-control AI proxies at a lower cost
TechCrunch
1h ago
Ai Focus
Explainable AI startup Goodfire launches a cheaper AI monitoring solution: it doesn't just look at model outputs but reads internal signals during the model's operation. This tool is already available to Baseten customers, and Goodfire claims that in Kimi K3 tests, its cost is much lower than that of the step-by-step review AI models, and it can capture most malicious hacker sessions.
Helpful
No.Help

Usually, the approach to ensure that AI behaves properly is to have another AI monitor it closely. This has always been the default solution, but when the agent runs continuously for several hours and processes a volume of text equivalent to several novels, the costs can rise rapidly.

A startup focused on explainability – that is, on understanding how the AI model works internally – Goodfire launched a cheaper solution on Thursday: monitors that directly observe what happens inside the AI model while it is working, rather than just reading what it writes. These monitors are now available to customers of Baseten. Baseten hosts and operates AI models for other companies.

Baseten's Base Labs announced last month the establishment of a security cooperation relationship with Goodfire and AI platforms Hugging Face.

Prior to this release, there has been a series of incidents this year where AI proxies managed to escape from the testing environment, including one where a OpenAI proxy breached the protections of Hugging Face. Goodfire, which is built around its first monitor and is an open-source model, took advantage of a vulnerability in the sandbox this summer to access the internet as well as information on GitHub.

The system of Goodfire is somewhat similar to airport security checks. Small detectors, known as probes, read the internal signals of the model at every step of the proxy's operation, just like security gates check each passenger. Only when a probe emits an alarm will another AI model intervene, similar to a further manual bag inspection.

Customers of Baseten can choose which risks to monitor, including aggressive hacking activities, the misuse of chemical and biological weapons, as well as rewarding hacking behaviors. They can also decide on the automatic response methods: recording incidents, submitting them for manual review, or directly rejecting requests.

Goodfire indicates that the operating cost of this method is also lower. Most AI monitors are independent models, which require re-reading everything that the monitored model has done, thus increasing time and cost. The probes of Goodfire, on the other hand, take advantage of the calculations that are already being performed by the model during runtime.

Goodfire CEO Eric Ho said last week at the MAD Podcast of venture capitalist Matt Turck: “Internal activation monitoring is really cheap because it reuses the computations from forward propagation. So the model is already doing this token calculation. All you do is take the already computed intermediate neural activations and then run a classifier on top of those internal computations.” In short, the model is already performing these mathematical operations, and the probe just reads the results.

Goodfire states that in the testing of Kimi K3, the cost of monitoring approximately 1,500 sessions was about $51; in contrast, a cheaper AI model that checks each step gradually costs about $233, while the top-tier model costs around $10,000. The probe captured 94% of malicious hacker sessions and sent 8.7% of harmless sessions for further inspection.

The company stated that running four probes simultaneously would only increase the time required for the model to start responding by less than 2%.

Goodfire, the Chief Technology Officer and Co-Founder, said: "The biggest advantage is that you can catch things before they happen. We can detect during the evaluation or training phase if a model might be susceptible to hacking attacks."

This solution is mainly aimed at open-source models. Developers can download these models and remove their security protections, and unlike closed-source laboratories, they do not come with monitoring capabilities built-in for running on their own systems.

Balsam said, "The damage that can be caused by individuals using open-source models is much smaller compared to using computing clusters like inference service providers, which bear most of the responsibility. When we reach the moment of open-source Mythos, it will become clear that models need to have safeguards deployed during inference."

Recent research, including that of Kimi K3 and the leading open-source models such as GLM 5.2, has found that between 50% and 96% of runs in AI proxy tests exhibited rewarding hacker behavior.

Goodfire is not the first company to try this approach. Google DeepMind stated in January that its research provided a basis for the deployment of misusage detection probes in Gemini.

Balsam said that these monitors are the near-term results of a longer-term research goal: to reverse-engineer a LLM so that its behavior can be traced back to where it appeared during training. "We want to turn the magic of training models into precision engineering," he said.

Tip
$0
Like
0
Save
0
Views 26
HKWDB reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Trump says the U.S. will not attack Iran before midterms
Trump said on Thursday that the United States would not attack Iran before the midterms on November 3, and described discussions with Tehran as "fructuous." Previously, he had stated that the government was considering launching a new round of strikes against Iran before the elections, which temporarily pushed up oil prices.
CNBC
·2026-10-09 01:01:56
5
Netflix releases trailer for new series "Altruists," focusing on scandals involving Sam Bankman - Fried, Caroline Ellison and FTX
Netflix releases the first trailer for the new series "Altruists." This eight-episode series will be available on November 19th and tells the story of a massive fraud case involving billions of dollars that led to the collapse of the cryptocurrency empire FTX. In the series, Anthony Boyle plays Bankman-Fried, while Julia Garner portrays Caroline Ellison, a former executive in the cryptocurrency industry who is also her ex-girlfriend.
Businessinsider
·2026-10-09 01:01:52
6
JULES chooses Ekinops network security solutions to enhance web page security and align with the changes in digital usage trends
JULES, headquartered in Lubei and employing 2,800 staff, has chosen the security service edge solution from Ekinops to enhance internet usage visibility, simplify filter policy management, and strengthen employees' protection against web page risks. This solution consists of SWG, local DNS, and CASB components, covering both headquarters and branch office environments.
PR Newswire
·2026-10-09 00:51:07
11
Arctiq has been rated as CRN Triple Crown Award for the third consecutive year
Arctiq announced that CRN has been recognized as the winner of the 2026 Triple Crown Award award, marking the third consecutive year that the company has received this honor. This award is given to solution providers that have made it onto three important lists of CRN.
PR Newswire
·2026-10-09 00:51:03
11
The U.S. Army awards a $14.5 million contract to webAI to deploy AI to soldiers' equipment
The U.S. Army awarded webAI a fixed-price contract worth $14.5 million through Project ARIA, requiring them to deploy AI that can operate in cloud-free, interfered-with, and intermittent network environments onto the devices already carried by soldiers. webAI stated that their software will run directly on the military's existing laptops, handheld devices, and vehicle-mounted computing devices, and will connect the various nodes through their Intelligence Delivery Network.
PR Newswire
·2026-10-09 00:41:36
13
View More