ASUS Next-Gen AI Servers Overcome Agentic AI Bottlenecks

HPC Supercomputing IT Infrastructure AI

As agentic AI moves from proof-of-concept to full-scale production, enterprises evaluating AI server options increasingly need to match server platforms to the scale and workload profile of their AI agent deployments. To address this, ASUS has launched the ASUS Vera C2 2U MGX model, XA P2N-E2, built on the NVIDIA® Vera CPU, alongside the ASUS RS700A-E14B, RS720A-E14B, and XA P8A-E14A series based on the AMD EPYC 9006 processors. Together, these servers help enterprises build high-performance Agentic AI infrastructure capable of supporting both complex task execution and large-scale AI agent deployment, accelerating enterprise adoption of agentic AI. 

 

Historically, AI servers have been optimized primarily around GPU compute. But as enterprise AI shifts toward agentic AI, workloads increasingly revolve around a continuous ‘agent loop’ — task planning, code execution, external tool invocation, and reflection/correction — that is heavily CPU-dependent. As the CPU takes more of this critical execution work, conventional AI server architectures are becoming a bottleneck for agentic AI. To meet the demands of mission-critical agentic AI tasks and large-scale deployment, next-generation server platforms must address four key hardware bottlenecks: 

 
  • Insufficient single-thread performance at scale: The speed at which a single AI agent completes a task depends heavily on CPU single-thread performance. When an agent executes Python scripts, compiles code, or calls external APIs within a secure sandbox — all serial workloads — weak single-thread performance slows processing and can leave GPUs idle waiting on the CPU. Next-generation CPUs must therefore deliver exceptionally high, stable, and predictable single-thread performance, and just as importantly, sustain that performance across every core under full-socket load, since large-scale agentic AI workloads depend on consistent per-core throughput even when all cores are active. 

  • Inadequate per-core memory bandwidth: Agentic AI typically involves many AI agents processing tasks concurrently. When a large number of CPU cores simultaneously handle code execution, tool calls, or RAG vector retrieval, insufficient memory bandwidth can turn memory channels into a bottleneck, reducing the bandwidth available to each core. Solving this requires a next-generation memory architecture with multiple channels and high bandwidth, ensuring every core retains adequate memory bandwidth even under full load — improving overall data access efficiency and processing performance. 

  • Cross-die data access latency: For latency-sensitive, mission-critical workloads such as real-time transaction monitoring or dynamic supply chain orchestration, a multi-die CPU design can introduce added latency as data moves between dies, increasing an AI agent's response time. This effect can compound further when multiple agents operate concurrently or collaborate with one another. Addressing this requires next-generation CPUs with highly coherent, ultra-low-latency inter-die interconnects — or a single-die architecture altogether — to minimize the wait time caused by cross-die data transfers and improve the real-time responsiveness and stability of multi-agent collaboration. 

  • Lack of hardware-level sandbox isolation: When AI agents are authorized to execute high-risk actions such as trades, order placement, or access to sensitive data, the absence of hardware-level security isolation means that a prompt injection attack or other malicious activity could extend its impact into core enterprise systems. Strengthening this defense requires next-generation AI servers with hardware-level sandboxing that isolates the underlying execution environment, establishing independent security boundaries for each AI agent and reducing the risk of unauthorized access or data leakage. 

 

To address these agentic AI hardware requirements, ASUS has introduced a new lineup of AI servers — the ASUS XA P2N-E2 and the ASUS RS700A-E14B, RS720A-E14B and XA P8A-E14A series — engineered across single-thread performance, memory bandwidth, real-time data access, sandbox isolation, and expansion capability to comprehensively support enterprise mission-critical workloads and large-scale AI agent deployment. 

 

XA P2N-E2: Extreme Single-Thread Performance for Complex AI Agent Tasks 

 

XA P2N-E2 is built around the NVIDIA Vera CPU, purpose-designed for agentic AI. With 88 Olympus cores, the Vera CPU sustains stable, consistent performance across every core even under full-socket load, delivering outstanding single-thread performance for large-scale AI agent workloads — up to 1.8x that of traditional x86 CPUs. 

 

The server also provides hardware-level sandbox isolation to strengthen the security of agent code execution. Backed by the Vera CPU's up to 1.2TB/s of LPDDR5X memory bandwidth and 3.4TB/s of high-speed internal interconnect bandwidth, the system accelerates code execution within agent sandboxes to 1.8x that of traditional x86 infrastructure and speeds up coding workflows by up to 1.5x — well suited to workloads involving frequent tool calls and sandboxed code execution. 

 

XA P2N-E2 is also ideal for agentic AI applications involving highly complex individual tasks, and supports up to two NVIDIA RTX PRO 6000 Server Edition GPUs to accelerate complex model inference for enterprise AI agents. The system is built on NVIDIA's MGX modular architecture, integrating CPU, GPU, and software stack to optimize overall system performance and reduce integration overhead. 

 

ASUS RS700A, RS720A and XA P8A-E14A Series: High Core Density for Massive Agent Fleets and Frequent RAG Retrieval 

 

Built on the AMD EPYC 9006 processors, the ASUS RS700A-E14B, RS720A-E14B and XA P8A-E14A series are purpose-suited for running large numbers of AI agents with frequent RAG retrieval. With up to 256 cores per CPU socket — an extreme level of core density — the platform delivers roughly 70% higher overall performance than the previous generation. This high core density allows each AI agent to be allocated a dedicated execution thread, significantly improving task execution efficiency and enabling thousands of agents to run tasks in parallel, meeting the demands of large-scale, high-density deployments. 

ASUS RS700A-E14B series


On the memory and data-transfer side, the AMD EPYC 9006 servers platform supports up to 16-channel DDR5 memory, delivering as much as approximately 1.6TB/s of memory bandwidth — more than double that of the previous generation — ensuring every core retains sufficient bandwidth even as large numbers of cores access data simultaneously. Combined with the high-speed data transfer capability of PCIe 6.0, this substantially reduces I/O latency when AI agents frequently query RAG vector databases, improving overall RAG performance. 


ASUS RS720A-E14B Series

 

Overall, the ASUS RS700A-E14B, RS720A-E14B and XA P8A-E14A series are best suited for enterprises that prioritize compatibility with existing x86 software, support for complex workloads, and deployment flexibility. ASUS offers 1U, 2U, and 6U form factors, allowing enterprises to choose based on the number of AI agents, data center space, GPU expansion needs, and compute requirements — balancing high-density computing with room for future growth. 

 

With high-performance CPUs, high-bandwidth memory, low-latency architecture, hardware-level security, and flexible expansion design, the ASUS agentic AI portfolio addresses the compute, security, and deployment needs of diverse agentic AI workloads — helping enterprises deploy AI agent applications reliably and scale them with confidence. 

 
  • Service and warranty coverage may depend on country and territory. Service may not be available in all markets. We recommend that you check with your local retailers to confirm the options available.
  • Must be purchased and activated within 90 days of your ASUS product purchase date via ASUS Premium Care.