Nvidia is making a new move in the artificial intelligence industry. Instead of only creating bigger and more powerful AI models, the company is working on models that can complete tasks faster and at a lower cost. Nvidia has released a new AI model called Nemotron 3.5 Lightning, which is designed to handle simple, repetitive tasks that are often performed by AI agents. Nvidia is also making the model available to developers and businesses so they can use and customize it.
AI agents are programs that can do more than answer questions. They can complete tasks on their own, such as researching a topic, writing computer code, searching a database, checking their work and organizing information without a person guiding them through every step. As AI agents become more common, companies need models that can handle many tasks quickly because using a very large and expensive AI model for every small job can take more time and cost more money.
That is where Nvidia says its new model can help. Nemotron 3.5 Lightning has 30 billion total parameters, but only about 3 billion are used at one time. It uses a system called a “mixture of experts,” which allows the model to use only the parts it needs for a specific task. A simple way to think about it is to imagine a school with hundreds of students. If a teacher has a math question, there is no reason to ask every student to solve it. The teacher can choose the students who are best at math, and Nvidia’s model works in a similar way by using the parts of the system that are best suited for each task.
Nvidia says this design allows Lightning to work much faster than some other models. The company says it can produce answers up to four times faster in certain situations. This speed could be especially useful for AI agents because an agent may need to complete many small tasks before finishing a larger job. For example, an AI agent helping a customer service worker might need to look up a customer’s account, find an order, check a delivery date and update information. These tasks may not require the most advanced AI model, so a smaller and faster model can handle the basic work while a more powerful model handles the difficult questions. This could help businesses save both time and money.
Nvidia is also introducing a technology called NeMo Switchyard, which is designed to help AI systems decide which model should handle each task. If an AI agent has a difficult problem, Switchyard could send it to a powerful AI model. If the next task is simple, it could send it to a smaller model like Nemotron 3.5 Lightning. This means companies would not have to use their most expensive AI model for every single task.
The idea is important because AI is starting to move beyond chatbots. Chatbots usually wait for a person to ask a question before responding, while AI agents can take several steps on their own to complete a goal. A company could use an AI agent to help prepare a report, for example, with the agent searching for information, organizing the data and creating the report. If every step required a large AI model, the cost could quickly become high. Using smaller models for simple tasks could make these systems more affordable.
Nvidia is also making Nemotron 3.5 Lightning available for developers and businesses to use and customize. This gives companies more control over how they use the technology. Some businesses may want to run AI on their own computers or servers instead of sending information to another company, while others may want to change the model so it works better for a specific type of job.
That openness could become one of the most important parts of Nvidia’s strategy. Harsha Kumar, CEO of NewRocket, believes the same forces that made open-source software a foundation of the internet could eventually reshape enterprise AI. “If there was no open source code, there would be no internet,” Kumar says. As open AI models continue to close the performance gap with proprietary models, he believes businesses will have a harder time justifying why their AI infrastructure should be tied to only a small number of vendors.
For businesses, that could mean more choices. Instead of relying completely on a single AI company, companies could have the ability to choose, customize and even run different models depending on what they need. This could also create more competition between AI companies and potentially push down the cost of using AI.
The release is also important for Nvidia itself. The company is best known for making the computer chips that power many of today’s AI systems, but Nvidia is now developing more than just chips. It is also creating AI models and software. Making AI more efficient could actually help Nvidia sell more technology because if AI becomes cheaper and faster to use, businesses may decide to use it for more tasks. That could create even more demand for the computers and chips needed to run AI.
Nemotron 3.5 Lightning is not meant to replace the most powerful AI models. Instead, Nvidia is promoting a future where different AI models work together. One model could handle difficult questions, another could write code, another could search for information and a smaller model could handle simple, repetitive jobs. This could make AI systems more like a team of workers, with each member handling the tasks they are best at.
Nvidia’s new model shows that the future of AI may not be only about creating the biggest and smartest models. It may also be about creating AI that is fast, affordable and open enough for businesses to control. As companies begin using AI agents for everyday tasks, smaller open models like Nemotron 3.5 Lightning could become an important part of how AI works. The goal is simple: use the right AI model for the right job instead of using the most powerful model for everything.






























