Proof for AI products. AI quality, verified.
Plune is quality infrastructure for teams shipping with LLMs — methodological rigor, human-in-the-loop, and CI/CD quality gates. Ship AI with confidence.
Early, but real — already on npm and running in CI. Open-source.
Cairn generates tests across your app's surfaces. Plune evaluates LLM behavior and gates releases in CI. Two separate layers — use either, or both.
An autonomous QA agent that explores your app and leaves a trail of tests.
- Logs in with a saved Playwright session, explores each page (ARIA snapshot + screenshot), and verifies every locator before trusting it.
- Writes methodology-grade test cases (ISO/IEC/IEEE 29119-4), then generates POM-style
@playwright/test. - Self-validates, self-repairs (keep-best), and self-improves via Langfuse — optional, so it runs fully offline.
- Decides what to automate (ATC) vs. leave to a human (MTC); modes
design·automate·explore; interactive TUI. - Surfaces today: UI. Next: API, unit, docs.
@plune-ai/cairn· Apache-2.0 · TypeScript · Node 20+ · formerly Lex-Bot, relicensed GPL-3.0 → Apache-2.0
Assertion-testing for LLM behavior, with regression gates in CI.
@plune-ai/cliruns an assertion suite against your provider (Anthropic · OpenAI · OpenRouter) and returns a pass/fail report — locally, in CI, or as a regression diff between runs.- 10 assertion types — from exact-match and
json-schematollm-judge, semantic-similarity, and RAG metrics (faithfulness, answer-relevance, context-precision). - Commands:
plune run·plune report·plune diff·plune init. eval-actionwraps the CLI: on every PR it runs your evals, diffs against the base branch, and leaves one sticky comment (what regressed, what improved) — optionally blocking merge on a pass→fail regression.@plune-ai/cli· MIT ·eval-action· MIT
- v0.1 — CLI with assertions (exact-match, contains, json-schema, llm-judge) — shipped
- v0.2 — GitHub Action + sticky PR comments — shipped
- v0.3 — Cloud sync + dashboard
Alongside the roadmap, Cairn opens a new pillar — test generation (UI today; API, unit, and docs next).
- cairn — autonomous agent that explores your app and generates Playwright tests.
- cli — assertion test-runner for LLM behavior, with CI regression diffs.
- eval-action — GitHub Action that runs Plune evals and gates PRs.
Building — or testing — AI products? Say hello: hello@plune.ai.
Building from Ukraine 🇺🇦