How NVIDIA's AI factory platform balances maximum performance and minimum latency, optimizing AI inference to power the next industrial revolution. When we prompt generative AI to answer a question or create an image, large language models generate tokens of intelligence that combine to provide the result. One prompt. One set of tokens for the answer. This is called AI inference . Agentic AI uses reasoning to complete tasks. AI agents aren't just providing one-shot answers. They break tasks down into a series of steps, each one a different inference technique. One prompt. Many sets of tokens to complete the job.
The engines of AI inference are called AI factories—massive infrastructures that serve AI to millions of users at once. AI factories generate AI tokens. Their product is intelligence. In the AI era, this intelligence grows revenue and profits. Growing revenue over time depends on how efficient the AI factory can be as it scales. AI factories are the machines of the next industrial revolution.