</>MCP Agents Market
Plugin

Model-Optimizer

by NVIDIA3.7kPythonUpdated 2026-09-01

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Claude CodeCodex

Model-Optimizer is a Claude Code plugin from NVIDIA that provides access to state-of-the-art deep learning model optimization techniques including quantization, pruning, neural architecture search, distillation, and speculative decoding. It compresses PyTorch, Hugging Face, and ONNX models for deployment to frameworks like TensorRT-LLM, vLLM, and SGLang, enabling faster inference with reduced memory footprint. The plugin integrates NVIDIA's ModelOpt library directly into AI-assisted development workflows, allowing developers to optimize large language models and diffusion models through natural language conversations.

Key Features

Post-training quantization (PTQ) supporting FP8, NVFP4, INT8, and INT4 formats for 2x-4x compression
Quantization-aware training (QAT) and distillation to recover accuracy after aggressive quantization
Model pruning and Minitron techniques to reduce parameter count and memory footprint
Speculative decoding to train draft modules for predicting multiple tokens during inference
Support for LLMs, VLMs, and diffusion models from Hugging Face, PyTorch, and ONNX
Direct export to TensorRT-LLM, vLLM, SGLang, and TensorRT deployment frameworks
Integration with Megatron-LM, Megatron-Bridge, and Hugging Face Accelerate for training workflows
Pre-quantized checkpoints available on Hugging Face for immediate deployment

Use Cases

  • 01Quantizing large language models like Llama, DeepSeek-R1, or Nemotron to FP8/FP4 for faster inference
  • 02Compressing vision-language models (VLMs) for deployment on resource-constrained hardware
  • 03Pruning and distilling models to create smaller, faster variants while maintaining accuracy
  • 04Optimizing Stable Diffusion and other diffusers models for near 2x speedup with INT8 quantization
  • 05Preparing models for production deployment in TensorRT-LLM, vLLM, or SGLang inference engines
  • 06Creating speculative decoding draft models to reduce latency in autoregressive generation

Related Plugins

View more

Model-Optimizer — FAQ

What is the NVIDIA Model-Optimizer Claude Code plugin?+

It's a Claude Code plugin that integrates NVIDIA's Model Optimizer library, providing AI-assisted access to model compression and optimization techniques like quantization, pruning, and distillation for deep learning models. The plugin enables developers to optimize LLMs, VLMs, and diffusion models through conversational interactions within Claude Code.

How do I install the Model-Optimizer plugin in Claude Code?+

Run 'claude plugin marketplace add https://github.com/NVIDIA/Model-Optimizer.git' to add the marketplace, then 'claude plugin install modelopt@modelopt' to install the plugin. The plugin will then be available in your Claude Code workspace.

Which AI clients support the Model-Optimizer plugin?+

The plugin is designed for Claude Code. The README also mentions Codex support with separate installation commands. It does not work with Claude Desktop, ChatGPT, or other AI assistants.

Do I need API keys or NVIDIA credentials to use this plugin?+

The README does not mention API key requirements for the plugin itself. However, you'll need a compatible NVIDIA GPU and CUDA environment to actually run the model optimization workloads, along with the nvidia-modelopt package installed via pip.

Is the Model-Optimizer plugin free to use?+

Yes, Model-Optimizer is open source under the Apache 2.0 license and free to use. The plugin provides access to the library's capabilities within Claude Code at no cost.

What model formats does Model-Optimizer support?+

It supports models from Hugging Face Transformers and Diffusers, PyTorch native models, and ONNX models. Optimized checkpoints can be exported for deployment in TensorRT-LLM, TensorRT, vLLM, and SGLang.

How do I install Model-Optimizer?+

Open the source repository on GitHub and follow its README. Model-Optimizer is a plugin — MCP Agents Market links you directly to the official repo.

Is Model-Optimizer free?+

Model-Optimizer is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches