Full speed ahead: Despite calls to slow AI down, its support structure is on the fast track
AI-summarised brief · reviewed before publication
At the AI Infra Summit in Santa Clara, senior engineers from Amazon Web Services, Oracle, Broadcom, Qualcomm and d‑Matrix detailed rapid advances in hardware and software designed to meet soaring artificial‑intelligence demand. Speakers highlighted the shift from GPU‑heavy training workloads in 2025 to inference‑centric processing in 2026, stressing that token consumption and power use have become critical cost drivers. Qualcomm’s Tony Pialis warned that “tokens per watt” is the new battlefield metric, while AWS senior vice president Peter DeSantis promoted the Graviton 5 Arm‑based CPU as the cornerstone for energy‑efficient inference at scale. Participants also flagged memory bottlenecks, noting that a typical AI server now requires roughly eight times more RAM than a conventional server, prompting a race to develop higher‑density, lower‑latency memory solutions.
💡 Why It Matters
- · The push to overhaul AI infrastructure now determines whether exploding inference workloads become sustainable or cripple enterprise adoption.