</>MCP Agents Market
Skill

hf-mem

by alvarobartt939PythonUpdated 2026-09-07

A CLI to estimate inference memory requirements for Hugging Face models, written in Python.

Claude Code

hf-mem is a Python-based CLI tool that estimates inference memory requirements for models hosted on the Hugging Face Hub, and can be added as an agent skill for AI coding assistants. It calculates memory footprints for Transformers, Diffusers, and Sentence Transformers models by reading Safetensors and GGUF metadata via HTTP range requests, without downloading full model weights. Developers can run it as a standalone command-line utility or integrate it as a skill that enables AI agents to automatically estimate model memory needs during code generation workflows.

Key Features

Estimates inference memory requirements for Hugging Face models including Transformers, Diffusers, and Sentence Transformers
Supports both Safetensors and GGUF model weight formats with metadata-only HTTP range requests
Experimental KV cache memory estimation for large language models and vision-language models with configurable batch size and sequence length
Breakdown of Mixture of Experts (MoE) models showing base model and expert weights separately
CLI and programmatic Python API with both synchronous and asynchronous interfaces
Lightweight dependency footprint requiring only httpx2
Integration as Hugging Face Hub CLI extension (hf mem) and agent skill for coding assistants
Per-file memory estimation for repositories containing multiple GGUF quantized variants

Use Cases

  • 01Determining hardware requirements before deploying Hugging Face models to production
  • 02Comparing memory footprints across different GGUF quantization levels for the same model
  • 03Enabling AI coding agents to automatically check if a model fits available GPU memory during development
  • 04Estimating total memory including KV cache for inference workloads with specific context lengths and batch sizes
  • 05Quickly assessing memory needs for diffusion models or embedding models without downloading weights
  • 06Auditing memory requirements across multiple model variants in a Hugging Face repository

Related Skills

View more

hf-mem — FAQ

What is hf-mem?+

hf-mem is a command-line tool and agent skill that estimates how much memory is needed to run inference with models from the Hugging Face Hub. It works by reading model metadata from Safetensors or GGUF files without downloading the full weights.

How do I install hf-mem as an agent skill?+

Place the SKILL.md file in the skills directory where your coding agent looks for capabilities, such as .claude/skills/hf-mem/SKILL.md. The agent will then be able to discover and use hf-mem when it needs to estimate model memory requirements.

Which AI clients work with hf-mem as an agent skill?+

The README mentions compatibility with coding agents that support the agent skills format introduced by Anthropic, such as Claude-based development assistants that read SKILL.md files from designated skill directories.

Do I need API keys to use hf-mem?+

No API keys are required for public Hugging Face models. The tool uses HTTP range requests to read model metadata directly from the Hugging Face Hub without authentication for publicly accessible repositories.

Is hf-mem free to use?+

Yes, hf-mem is an open-source tool available on GitHub. However, it is still experimental (pre-v1.0.0) and subject to breaking changes across releases.

How do I run hf-mem from the command line?+

The recommended way is using uvx with the command 'uvx hf-mem --model-id MODEL_NAME' where MODEL_NAME is any Hugging Face model identifier. You can also install it as a Hugging Face CLI extension and run 'hf mem' commands.

How do I install hf-mem?+

Open the source repository on GitHub and follow its README. hf-mem is a skill — MCP Agents Market links you directly to the official repo.

Is hf-mem free?+

hf-mem is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.

Related searches