Industry story
GPT-5.4 and NVIDIA Nemotron 120B land in AWS GovCloud via Kiro
ai-in-adtech cloud-costs engineering
GovCloud getting GPT-5.4 and NVIDIA Nemotron 120B through Kiro isn't a checkbox move — it's AWS signaling that agentic workloads are now a serious federal procurement target. GPT-5.4 brings a 272K context window and multi-step reasoning to classified-adjacent infrastructure; Nemotron's hybrid architecture activates only 12B of 120B parameters and runs at a 0.25x credit multiplier, making it the cost-efficient workhorse for agencies that need volume. The tension to watch: whether sovereign-cloud demand is large enough to justify the operational overhead of isolated queues and separate model pipelines, or whether this is AWS planting a flag ahead of actual budget.
Full analysis
AWS has added two new models to Kiro — its AI-powered IDE and CLI — specifically for the AWS GovCloud (US-West) Region, which serves U.S. government and regulated workloads. OpenAI's GPT-5.4 arrives with a 272K context window and is positioned for multi-step agentic workflows (sequences of AI actions that use tools, interpret context, and verify outputs across multiple steps), running on Bedrock's next-generation inference engine with isolated queues for resilience. NVIDIA's Nemotron 3 Super 120B, an open-weight hybrid mixture-of-experts model (an architecture that routes inputs to a subset of specialized sub-networks, activating only 12B of 120B parameters here for compute efficiency), offers a 256K context window and is priced at a 0.25x credit multiplier — making it a cost-efficient option for agentic tasks in sovereign cloud environments.
Comments