marin
Open-source framework for the research and development of foundation models.
Marin is an open-source framework providing agent skills for foundation model research and development, particularly large language model training. The platform offers reusable agent skills for tasks like adding scaling heuristics, datasets, and managing complex ML workflows. Developed by Stanford CRFM and Open Athena, it includes documented skills stored in .agents/skills/ directories that guide AI assistants through processes such as data curation, tokenization, pretraining, and evaluation. The framework supports training at massive scale, from tiny models to frontier mixture-of-experts systems with 500+ billion parameters.
Key Features
Use Cases
- 01Training large language models from scratch with documented, reproducible workflows
- 02Building mixture-of-experts models with load balancing at scale
- 03Creating scaling laws to predict large model performance from smaller experiments
- 04Curating and preprocessing training datasets with tokenization pipelines
- 05Extending foundation model training to non-text domains like audio, DNA, or proteins
- 06Learning LLM development best practices through documented experiments and agent-guided workflows
Related Skills
View moresuperpowers
An agentic skills framework & software development methodology that works.
skills
Skills for Real Engineers. Straight from my .agents directory.
skills
Public repository for Agent Skills
ponytail
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
marin — FAQ
What is Marin?+
Marin is an open-source framework for foundation model research that includes agent skills to guide AI assistants through complex ML workflows like data curation, model training, and evaluation. It provides reusable skills in .agents/skills/ directories alongside a complete training platform.
How do I install Marin?+
Install Marin by following the installation guide in the documentation at docs/tutorials/installation.md. The framework can be used as a library for experiments or accessed through its agent skills for guided workflows.
Which AI clients work with Marin agent skills?+
Marin agent skills are stored in both .agents/skills/ and .claude/skills/ directories, indicating compatibility with Claude and other agents that support skill loading from these standard paths.
What prerequisites are needed to use Marin?+
You'll need Python for the core framework, and access to compute resources (CPUs for small experiments, GPUs or TPUs for larger training runs). No API keys are required as Marin is fully open-source.
Is Marin free to use?+
Yes, Marin is completely free and open-source software. All code, checkpoints, datasets, and agent skills are publicly available on GitHub and Hugging Face.
What types of models can I train with Marin?+
Marin primarily focuses on large language models but has been successfully used for audio-text models, DNA models, and protein models. The framework scales from tiny experimental models to frontier mixture-of-experts systems with 500+ billion parameters.
How do I install marin?+
Open the source repository on GitHub and follow its README. marin is a skill — MCP Agents Market links you directly to the official repo.
Is marin free?+
marin is an open-source project hosted on GitHub. Check the repository for its license and any usage requirements.