As the evolution speed of models exceeds that of chip cycles, chip architects must strike a balance between flexible computing power, data mobility, and advanced defense measures for security.
Key points:
The rapid development of edge technologies is forcing chip architects and design teams to rethink how to integrate AI into vertical industry applications. The fundamental challenge lies in the mismatch between timelines, as engineering teams must design for AI models, workloads, and deployment requirements, which often continue to change after the silicon chip design has been finalized.
In addition, edge performance in the real world rarely depends on NPU peak TOPS. More often, the deciding factors are memory bandwidth, latency, power consumption budget, security, and data transfer. To keep up with evolving models, architects are increasingly relying on flexibility—by adopting heterogeneous computing, programmable data paths, scalable memory, secure on-the-fly updates, and close collaboration between software and hardware design.
But flexibility alone is not enough. AI introduces a wide range of attack surfaces in models, data, keys, firmware, and inference pipelines; therefore, robust security mechanisms must be designed into the hardware from the very beginning. At the same time, the fragmented toolchains across different runtimes and frameworks also make hardware-oriented optimizations as important as the raw computing power.
This structural change is redefining the boundaries of standard components.
Imagination Technologies Senior Director of Product Management, Rob Fisher, stated: " CPU is becoming more and more like GPU because they are trying to do more and more parallel processing. NPU is also becoming more and more like GPU as they strive to become more flexible, with CPU, NPU, and GPU acting as the convergence points in the middle. We find that an increasing number of customers are hedging for the future, rather than trying to extract maximum performance from today's systems. This means that we are no longer just looking at a GPU that can be inserted and only works in a GPU manner. This will have a system-level impact on the entire SoC."
When designing edge applications with AI as a product feature, chip architects and design teams need to consider many factors. Hardware choices will drive broader architectural decisions, but the challenges are far more than just NPU TOPS.
Infineon’s IoT, as well as the vice presidents and heads of PSOC Edge MCU and Edge AI Solutions in the Computing and Wireless Business Unit, Eduardo Montanez, stated: “The performance of the AI model depends on architectural details such as memory, latency, power consumption, and security.” This is also why the company’s edge microcontrollers include scalable memory, low-latency architectures, and energy-efficient designs to support battery-powered applications at the edge. They also incorporate security features designed to protect data and software IP.
Edge AI applications are already challenging chip architects and design teams to keep up with the rapid changes in models and workloads.
Cadence Director of Product Marketing Group, George Wall, stated: "As new networks continue to emerge, any solution must possess programmability to adapt to these new networks as well as the constantly changing activation functions and corresponding data types. At the same time, it must also meet stringent performance requirements."
This makes the collaborative design of software and hardware increasingly important. EDA of Siemens, AI, Group Product Manager Niranjan Sitapure, pointed out: "The AI model, NPU, memory hierarchy, as well as optimization techniques such as quantization, pruning, and distillation, must be optimized together to achieve efficient deployment."
The entire pipeline also must be optimized, as success depends on sensor integration, data preprocessing, inference, post-processing, and continuous model lifecycle management.
Ideally, chip architects and design teams should have a detailed understanding of what edge applications will be doing, including which models they will use, the required performance, available memory, and power consumption constraints, in order to design chips for the best PPA.
Rambus, Vice President of Silicon Intellectual Property Marketing, stated: "Unfortunately, that's not the case. Even after compression, the chip design cycle is still one year or longer ahead of the release of any product that contains that chip. During this period, it is reasonable for edge applications to evolve, including changes in underlying models, available memory, and power consumption. This evolution of applications forces chip designers to focus on creating chips that can adapt to these changing needs, which includes, but is not limited to, the efficient and secure transfer of data."
For others, AI acceleration is not the biggest challenge. What may limit the performance and efficiency of the chip could depend on the rest of the application.
Efficient Computer, the Chief Technology Officer and Co-Founder, stated: "However, architects and designers do not design it this way. They focus on neural network acceleration, claiming very high but purely theoretical TOPS, and there is no real application that can come close to achieving it. In fact, in many real workloads (such as sensor fusion, autonomous driving), the majority of the application still runs on a general-purpose but very inefficient core at SoC. The opportunity to truly bring intelligent applications to fruition is not about having faster NPU, but about making the other 80% of the application as efficient as that AI part."
Core Design Challenges
Performance does not depend solely on the computational power of the programmable processors used. Wall explains: “For example, the memory bandwidth must be sufficient to continuously supply data to the computing units, and the data needs to be located in a position that is easily accessible to the processors. Otherwise, the computing units will essentially be operating underutilized. It is difficult to determine the workload requirements, as chip designers may not know what networks consumers will be using in 3 or 5 years.”
This challenge is further exacerbated by the fact that edge AI is essentially a constrained-driven design problem. Karazuba states: "You are dealing with tight power consumption budgets, limited memory, mixed sensor inputs, and latency expectations that users can immediately perceive. The architecture must be efficient first before we can even talk about intelligence."
As the AI model evolves rapidly, the system architecture needs to have memory scalability to support large datasets and the AI model. At the same time, it is also necessary to reserve sufficient memory for the non-AI parts of the design.
Montanez states: "In addition, the design must support end-to-end secure configuration and deployment, and allow for on-site updates without compromising valuable data and software IP."
Moving model weights and activation values between DRAM and the computing engine can consume more power than the computation itself of AI. Sitapure states: "The evolution speed of AI models is faster than the silicon chip development cycle. New architectures, operators, and generative AI workloads continue to emerge, putting pressure on hardware platforms to maintain adaptability." He added, "For critical/high-risk applications, edge systems such as ADAS, industrial automation, and robotics require predictable latency, high reliability, and the ability to operate even under limited connectivity conditions."
While data movement may dominate performance and efficiency, other challenges can also be overlooked. Gobieski says, "The first problem is simple, but difficult to solve: intelligence must be built into the device. This means that a significant amount of engineering work is required before any application code can run, in order to compress the neural networks into the device's memory. The second problem has a standard answer, but that answer could backfire. Splitting the application across NPU, CPU, and DSP might seem like it takes advantage of each component's strengths, but it would increase development complexity due to the need for many different toolchains. Moreover, the resources saved would be spent on data movement between them."
The hardware team also faces uncompromising physical constraints, especially strict low-power consumption budgets. Secure-IC co-founder and Chief Technology Officer Sylvain Guilley pointed out that heat dissipation limitations and battery constraints determine computational efficiency, while mission-critical and security-critical edge deployments also require a high level of reliability. "To address these core challenges, it is necessary to move beyond proprietary and isolated hardware techniques and towards broader industry-level architectural standards collaboration. To establish standardized and open framework specifications for edge hardware, Cadence is advancing this work within the new OCP FCSA sub-project workflow."
Consider security from the very beginning.
Electronic developers who wish to incorporate AI into edge devices should integrate various security technologies from the very beginning of the system design process.
The starting point is the security platform, followed by secure applications. Synopsys Senior Director of Security Solution Product Management, Dana Neustadter, stated: "The security platform is the foundation. It establishes the base, but it does not automatically make the applications – or the AI outputs – credible. Developers still need to handle authorization, data protection, output verification, connection service security, and component assurance at the application layer. Assuming that the AI platform already provides a trust root, secure boot, authentication, encrypted communication, and a trusted execution environment, this lays the foundation, but it does not automatically make the applications – or the AI outputs – credible."
At the application layer, Neustadter indicates that developers must handle additional controls, including:
A common misconception is that as long as AI runs on a secure platform, the results of AI must be credible. Imagine a mobile robot whose model reports that the corridor is empty. If the software accepts this result without further verification, the robot might move and collide with unexpected obstacles. Neustadter explains: “AI is just one input for decision-making. It should not be the sole basis for the final decision. The application should assess the model’s confidence and perform additional checks before authorizing any action. It should also apply relevant rules and security strategies, and confirm the environment before the device takes action.”
Establishing a resilient platform does not guarantee privacy. Secure hardware alone cannot compensate for insecure data processing at the application layer. For example, smart cameras in shopping malls may analyze shoppers' behavior. Even if the devices themselves are secure, if the analysis results are stored in unencrypted databases, attackers may still have access to private customer information.
"Security hardware alone cannot compensate for the insecurity of data processing at the application layer. Developers should encrypt sensitive data and implement access control throughout its entire lifecycle. They should also retain data only when necessary and limit its use to authorized personnel and services," she said.
Models should also be authenticated in the same way as secure startups. Before loading a model, the system should verify its identity, integrity, signature, and origin. This is another reason why a strong trust root is crucial. Neustadter explains: 'Model confidentiality and access control should also be incorporated into the system's key management strategy.'
The edge AI system may also require additional keys. For example, third-party model providers may wish to prevent device owners or others from accessing the models without authorization. Therefore, model confidentiality and access control will become part of the overall key management strategy.
All models and components can introduce risks to the software supply chain. Developers often rely on open-source models, AI engines, plugins, and other external components. Models that contain hidden backdoors may generate errors or malicious output, and if the system treats them as trustworthy, secure hardware may still execute them.
Neustadter added, "Even if the system considers a compromised model to be trustworthy, secure hardware will still execute it. To reduce risks, developers should verify the source and integrity, validate suppliers, use trusted repositories, and maintain a software bill of materials."
When designing edge application security and mitigating the attack surfaces introduced by AI, engineering leadership should involve senior product management as early as possible. As indicated by Guilley of Cadence, this allows developers to conduct risk analysis and determine which areas at the hardware interface and software levels require mitigation measures.
He pointed out, "This early architectural foundation directly affects the cybersecurity resilience within the overall device strategy." He also stated, "Establishing this foundation allows us to view cybersecurity as an investment rather than a cost, enabling subsequent operations to be carried out securely 24/7 and transforming incidents from emergencies into part of a business continuity plan."
When engineering teams build complex network physical systems that must comply with cybersecurity, AI, and data privacy requirements, such a unified security and governance framework is particularly important.
Guilley said, "The design team must continuously deal with strict regulatory boundaries, especially those surrounding the EU's Cyber Resilience Act, the legal requirements for collecting personal data, as well as the overall legal obligations imposed by the broader EU AI Act."
It is also important to isolate the trust root from non-security code. Montanez states: "It is crucial to have a bootloader that manages code authentication. As AI migrates to the edge, most products still adopt a hybrid approach that combines with the cloud to support OTA or other communications. Therefore, edge AI applications may need to provide additional security protection for the connection links."
Overall, security must be anchored in hardware, from secure boot and trusted firmware to protected keys, memory, and model assets. AI increases the risk, as models, inputs, and outputs can all possess significant economic value.
Karazuba points out: 'It should be taken into consideration that economic value comes not only from the model itself or the data within the model, but also from the financial and legal implications of third-party tampering with this data.'
Introducing AI in edge devices will add capabilities and functions that were often absent in previous versions not based on AI. He said, "Given that security vulnerabilities can have greater financial and legal implications, device manufacturers must design edge devices to support AI with stronger security configurations."
Sitapure indicates that encryption is also a key element. Model weights, training data, and the inference pipeline should all be encrypted to prevent tampering, extraction, and unauthorized access.
Most importantly, security cannot be a matter of remediation after the fact. Wall says: "Malicious actors may attempt to attack the underlying model weights and implant their own models, thereby disrupting the operation of subsystems or exploiting other security vulnerabilities in the system. Fortunately, there are solutions ranging from local encryption and decryption engines to dedicated security islands."
More broadly speaking, the most dangerous failure in edge systems is not a crash, but rather the system continuing to operate while silently believing in an incorrect state. Real-time decision-making increasingly relies on the integrity of device-side reasoning, as edge systems are deployed in cameras, robots, vehicles, and industrial systems. These systems depend on NPU, GPU, and TPU to improve performance and efficiency, but they are usually not designed to ensure the integrity of reasoning under fault conditions.
In a recent white paper, security researchers from Keysight presented experimental diagrams and physical setups for fault injection and measurement. The results showed that precisely timed fault injections could disrupt the output of production-grade edge AI pipelines without modifying firmware, model weights, software configurations, or input data. On a commercial Rockchip RK3568 platform running YOLOv5s_ReLU, most faults had no observable impact or only changed the confidence scores; however, a few faults led to missed detections, false targets, and rare cases of high-confidence misclassifications, yet the system continued to operate normally. The effects of these faults could extend beyond a single inference error. When the compromised detection results were fed into downstream tracking systems, even brief or low-probability errors could be filtered out as transient noise, locked in as persistent tracking targets, or amplified into more frequent and longer-lasting erroneous system states.
In the observed sequences, repetitive and constrained failures transform brief detection errors into persistent error perceptions that persist even after the original failure window has ended. Once these errors are represented as system states, they can affect alarms, automated logic, and operator decisions, even though the platform appears to be healthy on the surface. This is of significant importance to the product team. Here, traditional measures of resilience, such as uptime or failure recovery, may not be sufficient for the AI drive system. Resilience must also include the integrity of reasoning, as well as the ability to prevent corrupted outputs from being considered credible system states.
At the same time, in some applications, the decentralization of computing also brings fundamental security advantages, as it reduces the extent to which sensitive information leaves the device. Gobieski says: "Running inference on the device means that sensitive data is not transmitted, so there is no communication channel that can be intercepted. Decentralized processing means that there is no central repository of sensitive data that can be compromised. The cost of compromising one device is much lower than that of compromising an entire fleet. The models themselves still pose the usual concerns regarding exposure, but these concerns are much smaller compared to keeping the data on the device."
Conclusion
To address the challenges of the AI landscape, engineering teams need to undergo a shift in mindset. Success no longer solely depends on maximizing theoretical peak computing power, but rather on achieving a balance between the lifespan of silicon chips and the rapid evolution of software. By prioritizing flexible architectures, collaborative design of hardware and software, and robust security measures from the very beginning, chip designers can build resilient platforms that are capable of adapting to any future AI workloads—even after the chip architecture has been finalized.
Related Reading
AI is rewriting the game rules of IP.
As the semiconductor ecosystem shifts towards AI, it is changing the way IP is created, verified, managed, and sold.
Chips are about to face security verification – but the standards may not be sufficient yet.
As AI, post-quantum cryptography, chiplets, and custom silicon expand their attack surfaces, semiconductor security is shifting from isolated protection to full-stack verification.











