Encord Active – Unit Testing for Computer Vision Models
Open-source toolkit to test, validate, and evaluate CV models. Find label errors, surface weak data, and prioritize labeling with 458 GitHub stars.
Tag
13 posts tagged #testing
Browse 13 posts tagged Testing, including practical setup notes, reviews, comparisons, and workflow patterns for engineers working with AI tools.
Open-source toolkit to test, validate, and evaluate CV models. Find label errors, surface weak data, and prioritize labeling with 458 GitHub stars.
A Playwright-style Rust library for driving native desktop apps on macOS, Windows, and Linux. MIT licensed.
Testronaut lets you write Playwright tests in plain English. Define missions, run real browsers, get AI-generated reports without brittle selectors.
TesterArmy is a YC P26-backed agentic QA platform. Describe test flows in plain English and let AI agents run end-to-end checks on web and mobile, with Slack alerts and CI integration.
MCPJam Inspector is a local dev client for testing MCP servers and ChatGPT apps. Chat with AI models, trace tool calls, validate OAuth flows, and run CI/CD evals.
Spec27 from Safe Intelligence is a spec-first testing platform for AI agents. It generates adversarial tests from declarative specs and validates vendor agents without SDK access.
guard-skills adds second-pass checks for AI-generated code, tests, docs, and WordPress or WooCommerce changes before merge.
Open-source MCP testing platform: Playground, OAuth conformance, evals across models, and a Skills tab that teaches agents how to use your MCP tools.
Facts turns project claims into checkable records for coding agents, with lifecycle tags and CLI verification. This guide covers install, workflow, commands, and where it fits in real teams.
Cekura is a testing platform for voice and chat AI agents with simulated conversations, LLM-judge and code-based evaluators, production monitoring, and CI integration.
Free in-browser tool to inspect, test, and chat with MCP servers across 30+ models, with a 10K-server registry and JSON tool-call inspector.
Canary reads your codebase, understands what a pull request changed, generates end-to-end tests for affected user flows, and runs them against preview apps.
BrowserBook is a Mac desktop IDE for writing, debugging, and replaying Playwright-based browser automations with deterministic test execution.