Testing
100 listings
superpowers
An agentic skills framework & software development methodology that works.
skills
Skills for Real Engineers. Straight from my .agents directory.
agent-skills
Production-grade engineering skills for AI coding agents.
playwright
Playwright is a framework for Web Testing and Automation. It allows testing Chromium, Firefox and WebKit with a single API.
puppeteer
JavaScript API for Chrome and Firefox
Front-End-Checklist
🗂 The essential checklist for modern web development, for humans and AI agents
openinterpreter
A coding agent for open models like Kimi K3
strix
Open-source AI penetration testing tool to find and fix your app’s vulnerabilities.
chrome-devtools-mcp
Chrome DevTools for coding agents
agent-browser
Browser automation CLI for AI agents
playwright-mcp
Playwright MCP server
oh-my-pi
⌥ AI Coding agent for the terminal — hash-anchored edits, optimized tool harness, LSP, Python, browser, subagents, and more
kilocode
Kilo is the all-in-one agentic engineering platform. Build, ship, and iterate faster with the most popular open source coding agent.
cua
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
deepeval
The LLM Evaluation Framework
react-doctor
Your agent writes bad React. This catches it
SeleniumBase
📊 APIs for web automation, testing, and bypassing bot-detection.
playwright-cli
CLI for common Playwright actions. Record and generate Playwright code, inspect selectors and take screenshots.
iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
ginkgo
A Modern Testing Framework for Go
lamda
Android Full-Stack Device Control Platform: WebRTC/H.264 remote desktop, UI/OCR/image-matching automation, one-click MITM, built-in Frida, proxy/VPN/frp/P2P networking, MCP/Agent, 160+ APIs, designed for multi-device clusters and engineered deployments.
vphone-cli
browser-tools-mcp
Monitor browser logs directly from Cursor and other MCP compatible IDEs.
reqable-app
Reqable issue track repo
XcodeBuildMCP
A Model Context Protocol (MCP) server and CLI that provides tools for agent use when working on iOS and macOS projects.
mobile-mcp
Model Context Protocol Server for Mobile Automation and Scraping (iOS, Android, Emulators, Simulators and Real Devices)
maildev
:mailbox: SMTP Server + Web Interface for viewing and testing emails during development.
autoresearch
Claude Autoresearch Skill — Autonomous goal-directed iteration for Claude Code. Inspired by Karpathy's autoresearch. Modify → Verify → Keep/Discard → Repeat forever.
darwin-skill
达尔文.skill —— 一个让你的Skill无限进化的系统:评估→改进→测试→保留或回滚 | Autoresearch-inspired autonomous skill optimization for Claude Code. Evaluate, improve, test, keep or revert.
ouroboros
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop. MCP server, 13 runtimes: Claude Code, Codex CLI, Gemini CLI, OpenCode, Copilot,
agents-cli
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
mcp-playwright
Playwright Model Context Protocol Server - Tool to automate Browsers and APIs in Claude Desktop, Cline, Cursor IDE and More 🔌
skills
Repository for skills to assist AI coding agents with .NET and C#
dalfox
🌙🦊 Dalfox is a powerful open-source XSS scanner and utility focused on automation.
unidbg
Allows you to emulate an Android native library, and an experimental iOS emulation
memlab
A framework for finding JavaScript memory leaks and analyzing heap snapshots
App
Welcome to New Expensify: a complete re-imagination of financial collaboration, centered around chat. Help us build the next generation of Expensify by sharing feedback and contributing to the code.
mockserver-monorepo
MockServer is an HTTP(S) mock server and proxy for testing that lets you mock APIs, inspect and modify live traffic, and inject failures. It supports HTTP/1.1, HTTP/2, gRPC, WebSockets, TCP and more on a single port, with additional support for HTTP/3, message brokers, and AI/LLM APIs.
Agentic-Bug-Hunter
AI-powered bug bounty hunting toolkit that works with or without subscription.
agent-device
Mobile app automation and verification for AI coding agents. CLI, MCP server, and typed Node.js API for iOS, Android, HarmonyOS, TV, web, macOS, and Linux.
httprunner
HttpRunner 是一款开源的 API/UI 测试框架,简单易用,功能强大,具有丰富的插件化机制和高度的可扩展能力。
anti-slop
Opinionated Oxlint rules for rejecting low-evidence TypeScript and JavaScript patterns
playwriter
Chrome extension & CLI to let agents control your browser. Runs Playwright snippets in a stateful sandbox. Available as CLI or MCP
skills
Browserbase's official collection of agent skills to access the web.
expect
Expect tests your agent's code in a real browser
langwatch
The platform for LLM evaluations and AI agent testing
roslynator
Roslynator is a set of code analysis tools for C#, powered by Roslyn.
grok-cli
An open-source coding agent for the Grok API
Devon
Devon: An open-source pair programmer
Papercut-SMTP
Papercut SMTP -- The Simple Desktop Email Server
VulnClaw
基于 AI Agent + MCP 工具链 + 渗透 Skill 编排, 配合大语言模型, 自然语言输入 → 自动完成「信息收集 → 漏洞发现 → 漏洞利用 → 报告生成」全流程。
cc-skills-golang
🧑🎨 A collection of Golang agentic skills that works
testsprite-cli
Official TestSprite CLI — AI-powered automated testing from your terminal
playwright-skill
General-purpose Playwright automation for coding agents
desloppify
Agent harness to make your slop code well-engineered and beautiful.
unlazy
Anti-laziness skill for AI agents. Core: the Depth Tree method, which splits a task N layers deep and gives every leaf the full time budget of the whole task, so effort multiplies with depth. Grounded in 2025-2026 research on model laziness, underthinking and premature completion.
llm-as-a-verifier
LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and medical agentic benchmarks.
vibium
The verification layer for coding agents
comet
Comet: agent skill harness for turning ideas into evaluated workflows
moemail
A cute temporary email service built with NextJS + Cloudflare technology stack 🎉 | 一个基于 NextJS + Cloudflare 技术栈构建的可爱临时邮箱服务🎉
web-quality-skills
Agent Skills for optimizing web quality based on Lighthouse and Core Web Vitals.
serve-sim
The `npx serve` of Apple Simulators.
pymobiledevice3
Pure python3 implementation for working with iDevices (iPhone, etc...).
Mano-P
Mano-P: Open-source GUI-VLA agent for edge devices. #1 on OSWorld (specialized, 58.2%). Runs locally on Apple M4 Mac mini/MacBook — no data leaves your device.Mano-P 是一个开源 GUI-VLA 项目,支持在 Mac mini/MacBook 上或通过算力棒本地运行推理,实现纯视觉驱动的跨平台 GUI 自动化操作。数据完全本地处理,支持复杂多步骤任务规划与执行。
commands
A collection of production-ready slash commands for Claude Code
argent
An agentic toolkit to control, debug, and profile iOS and Android apps. Made by Software Mansion.
greenlight
Pre-submission compliance scanner for the Apple App Store and Google Play. Scans code, privacy manifests, Android manifests, and IPA/APK/AAB binaries against the review guidelines. Offline, no account.
gomega
Ginkgo's Preferred Matcher Library
tdd-guard
Automated TDD enforcement for Claude Code
terraform-skill
Terraform & OpenTofu Skill for AI Agents - testing, modules, CI/CD, and production patterns
fable-method
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
harmonist
Portable AI agent orchestration with mechanical protocol enforcement. 186 agents, zero runtime dependencies.
smallcode
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
everything-claude-code-zh
everything-claude-code 中文翻译项目:完整的 Claude Code 配置集合(agents, skills, hooks, commands, rules, MCPs)。源自 Anthropic 黑客松获胜者的实战配置,助力中文工程师高效理解与使用 Claude Code。
babysitter
Babysitter enforces obedience on agentic workforces and enables them to manage extremely complex tasks and workflows through deterministic, hallucination-free self-orchestration
context-engineering-kit
Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source alternative.
pentest-ai
Open-source AI pentester that proves every finding. Machine oracles re-run each exploit; verified bugs ship a proof capsule you can replay yourself.
multi-agent-coding-system
Reached #13 on Stanford's Terminal Bench leaderboard. Orchestrator, explorer & coder agents working together with intelligent context sharing.
evo
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
moai-adk
Agentic development harness for Claude Code — SPEC-driven plan/run/sync, TRUST 5 quality gates, model+effort routing, and Claude×GLM multi-LLM cost control. Single Go binary, 16 languages, zero deps.
langchain-skills
SWE-AF
Autonomous software engineering fleet of AI agents for production-grade PRs on AgentField: plan, code, test, and ship.
cali
AI agent for building React Native apps
tsumiki
tsumiki
agent-md
Production-grade agent directives for autonomous coding agents: Claude Code, Codex, Cursor, Windsurf, Aider.
ux-ui-agent-skills
Turn Claude into a Senior Design Architect — DTCG design tokens, 42 components, WCAG 2.2 accessibility, any-framework code, 138 design systems, and runnable skills.
luban-skill
鲁班 | Luban — 把'能用的Skill'打磨成'能被装、能传播、能验证、能进化'的公共资产。Agent skill-polishing workshop: 验料·访行·过尺·慢刨·回炉
agent-qa
Open-source self-improving QA agent for software teams. A test harness with memory. Write tests in natural language for web and mobile. agent-qa learns from every run, adapts to UI changes, and catches regressions before you ship.
copilot-orchestra
Agents and workflow for GitHub Copilot
fablize
A Claude Code plugin that makes Opus behave like Fable — completion, evidence, and verification enforced as procedure. Ships only what a Fable-vs-Opus comparison proved transferable.
SkillForge
A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
skill-up
An evaluation and evolution tool for Agent Skills.
virtme-ng
Quickly build and run kernels inside a virtualized snapshot of your live system
arrakis
A fully customizable and self-hosted sandboxing solution for AI agent code execution and computer use. It features out-of-the-box support for backtracking, a simple REST API and Python SDK, automatic port forwarding, and secure MicroVM isolation. Perfect for safely running, testing, and backtracking
autoprompt-skill
Autoprompt is a coding-agent skill that cuts failures by 45% on agentic coding tasks.
pdd
Prompt Driven Development (PDD): The Last Programming Language™. Prompt files are source; code is generated output.
proofshot
Give AI coding agents eyes. Records browser sessions, captures screenshots, collects errors, and bundles proof artifacts — so humans can verify what the agent built.
polar
One-click Bitcoin Lightning networks for local app development & testing
shuru
A local-first microVM sandbox for running AI agents safely on macOS & Linux
android-mcp-server
An MCP server that provides control over Android devices via adb