The new SMITH framework enables small models to compete with larger models using less token, paving the way for scalable multi-agent collaboration.
Singapore, September 30th / PRNewswire / -- Appier (Tokyo Stock Exchange: 4180) announced today that its latest research paper " Joint Optimization of Tool Creation and Use for Large Language Model Agents " has been accepted by NeurIPS. NeurIPS is a top global conference on artificial intelligence and machine learning, often referred to as the " AI Olympics".
This paper proposes the SMITH (Schema-grounded Multi-task Iterative Tool Honing) framework, which is a reinforcement learning framework that enables AI to build tools within the same training cycle and effectively utilize these tools. Moreover, each tool is continuously optimized based on the results of solving real-world problems. This breakthrough addresses a key challenge in Agentic AI: while models can create tools, it is often difficult to make good use of them.
Research shows that small models trained using SMITH are capable of building reusable tools that can compete with those created by larger models, even on tasks that the models have never encountered before. These tools can also be shared between different models and tasks, significantly reducing the token costs associated with repeated inference. The paper was accepted by NeurIPS, highlighting Appier's research strength in improving the efficiency of AI agents and advancing multi-agent collaboration, as well as strengthening its position at the forefront of global Agentic AI research.
Dr. Zhou Xihan, co-founder and CEO of Appier, stated: "Humans transform the experience of solving problems into tools, so there is no need to start from scratch every time." AI agents are also evolving in the same way. This research demonstrates that agents can learn to build tools, continuously optimize them, and share proven tools across models of different scales, thereby making multi-agent collaboration more efficient and scalable. NeurIPS's acceptance of this paper further recognizes the forward-looking research and innovation of Appier. We will continue to bring Agentic AI into real applications and deliver measurable results for businesses."
AI Agents must be able to build tools, and also need to know whether the tools are effective.
As Agentic AI takes on more and more autonomous handling of enterprise workflows, selecting and invoking the correct external tools has become a key to successful deployment. However, many AI systems still rely on engineers to pre-build API or configure fixed tools, which must be rebuilt whenever there are changes in data sources, tasks, or business requirements. Even though AI is capable of creating tools on its own, the existing methods typically assign "creation" and "usage" to different models. As a result, the models that build these tools receive little feedback from real-world usage, and it is difficult to determine whether the tool descriptions are clear, whether the tools run reliably, or whether other models can invoke them correctly.
SMITH incorporates these two capabilities into the same training loop, allowing AI to learn how to build good tools and use them effectively at the same time. When the tool descriptions are vague, the parameter design is poor, or the operations fail, these outcomes are fed back into the model. The processes of building, using, verifying, and optimizing form a closed loop that continuously improves the quality of the tools. Research has found that training the creation and use of tools together is significantly more effective than training them separately.
Research scientist Lin Zheyan from Appier stated: "When SMITH uses tools to train its models, it only sees the descriptions and parameter specifications of those tools, not the underlying code. This means that whether each description is clear and whether the tools can be called correctly both receive direct feedback during the training process. Our experiments have also confirmed that other models can use these tools to solve problems more effectively. Looking to the future, we hope to build models that can continuously interact with the environment and undertake a wider range of tasks."
Learn from 4 examples and verify on 16 new problems.
SMITH adopts a training approach that progresses from easy to difficult. AI starts by learning a method from 4 simple examples and uses it to build a tool, then tests whether the model can generalize to 16 more difficult problems that have not been seen before. The system retains only those tools that can solve new problems and adds them to a shared tool library for use by multiple AI agents. As high-quality tools are continuously added, better tools will replace the weaker ones. Agents do not need to create tools for each task; instead, they can build on and reuse proven problem-solving capabilities.
This study presents three key findings:
- Small models can surpass large models in tool creation: A model with approximately 4 billion parameters, trained using SMITH, built tools that outperformed all other methods studied on unseen tasks. It even surpassed the baseline of a model with about 30 billion parameters when used to build tools immediately. Effective tool creation does not necessarily depend on larger models.
- Verified tools can be used across model scales: Tools built for small models also perform well when handling new tasks on lightweight models with only about 350 million parameters. These tools have also improved the performance of larger models, enabling agents to divide labor more flexibly and efficiently.
- Repetitive reasoning can be transformed into a reusable tool: SMITH converts repetitive reasoning into tools that can be called directly. In experiments, the average output decreased from 3206 token using traditional step-by-step reasoning to about 100 token, representing an efficiency improvement of approximately 32 times. When facing similar problems, AI does not need to execute the entire reasoning process every time, allowing for faster reasoning while maintaining task performance and reducing computational costs.
This research has opened up new directions for the company Agentic AI. Many daily operations are repetitive, such as converting financial indicators, processing data, querying reports, checking rules, and diverting customer service tickets. AI can transform these scattered methods into verified, shareable tools that can be called by any agent, thereby reducing the costs of repeated development and reasoning.
In the fields of advertising and marketing, agencies that handle customer data, personalization, customer service, and ad purchases can share proven tools and consistent business rules to work together more efficiently. Whether entering new markets, setting up new advertisers, or starting with limited data, companies can transform past successes into verifiable and scalable AI capabilities, helping agencies adapt to new tasks more quickly. Appier will continue to advance Agentic AI through forward-looking research, making AI the core engine for long-term business growth.
About Appier
Appier (Tokyo Stock Exchange: 4180) is a AI native Agentic AI as Service (AaaS) company that empowers business decisions with cutting-edge AdTech and MarTech solutions. Founded in 2012, the company's vision is to "make software intelligent and AI simpler." It is committed to helping enterprises transform AI into ROI through Ad Cloud, Personalization Cloud, and Data Cloud solutions. Currently, Appier has 17 offices in the Asia-Pacific, United States, and EMEA regions and is listed on the Tokyo Stock Exchange. For more company information, please visit www.appier.com; for more investor relations information, please visit ir.appier.com / en.










