For Claude Code & Codex users

How AI-native are you, really?

Vibemetric reads the AI sessions already on your Mac and scores how you actually work with AI: how you frame tasks, feed context, verify results, and steer your agents. Your score lives in the menu bar.

macOS 14+ · Apple Silicon and Intel · 1 MB download · free

Terminal
63
63/ 100AI NATIVE SCORESep 28
  • Task Framing & Specs4
  • Context Engineering7
  • Memory & Knowledge7
  • Tool Integration9
  • Verification8
  • Iteration Loops6
  • Orchestration7
  • Evals3
  • Human Steering8
  • Safety & Cost4

Next: Write the pass/fail checks before the AI starts

Open VibemetricScore me again

How it works

Scored from what you did, not what you say you do.

  1. 01 · about a second

    Scan

    Vibemetric reads the Claude Code and Codex transcripts already in ~/.claude and ~/.codex. It never writes to them.

  2. 02 · on your Mac

    Digest

    It condenses every session into a short timeline of what you asked, what the AI ran, and what failed, with keys, tokens and database URLs removed.

  3. 03 · about five minutes

    Judge

    Your own Claude Code reads the digest with read-only tools and applies the rubric. You get a score, a reason for each dimension, and one thing to change first.

The rubric

10 dimensions, 100 points.

Each dimension scores 0–10 against fixed behavioral anchors. Scores of 8 and up require the practice to carry into a later session. Installs, agent counts and PRs don’t earn points.

  • 01

    Task Framing & Specs

    Can your completion criteria accept or reject the result?

  • 02

    Context Engineering

    Does the AI get the smallest set of high-signal context it needs?

  • 03

    Memory & Knowledge

    Does saved knowledge change a decision in another session?

  • 04

    Tool Integration

    Do tool results and failures decide the next action?

  • 05

    Verification

    Is real behavior checked, not just “it compiles”?

  • 06

    Iteration Loops

    Does the same failure go through fix → rerun → stop?

  • 07

    Orchestration

    Are roles, handoffs, and the combined result verified?

  • 08

    Evals

    Do judged runs drive keep, revise, or retire decisions?

  • 09

    Human Steering

    Do your corrections become controls that prevent the next mistake?

  • 10

    Safety & Cost

    Does a policy connect risk, model, permissions, and budget?

STRONG · 8–10 DEVELOPING · 6–7 BOTTLENECK · 0–5

A sample card

Every score comes with the case that earned it.

The card explains each score through a real episode from your sessions, cites where to find it, and ends with one change to try on your next task.

This sample is fictional.

63/ 100

You give AI real tools it keeps reusing, and it checks live behavior on its own. But you rarely say what “done” means up front, so design tasks end in rounds of taste corrections.

  • Task Framing & Specs4/10
  • Context Engineering7/10
  • Memory & Knowledge7/10
  • Tool Integration9/10
  • Verification8/10
  • Iteration Loops6/10
  • Orchestration7/10
  • Evals3/10
  • Human Steering8/10
  • Safety & Cost4/10

Verification · 8/10

The AI proves behavior by driving the real site, then reruns the same check after the fix.

When a site header flickered on every page change, the AI wrote a small script that tags the header in the page, clicks a nav link, and checks whether the tag survives. It didn’t, which proved the header was being rebuilt.

After moving the header into the shared layout, the AI reran the same script. The header survived, and it checked the layout across pages before committing. A week later the same script gated another navigation change.

Claude Code · a41c9e02 · 2026-09-14 — header check failed → layout fix → same check passed

One improvement to make first

Write the pass/fail checks before the AI starts

“Before you change anything, restate these acceptance checks and how you’ll prove each one against the running app. Don’t report done unless every check has evidence.”

Privacy

It reads your AI transcripts. Here’s exactly what happens to them.

Vibemetric makes no network calls

The app has no server and no account. The only thing that talks to Anthropic is your own Claude Code, exactly as when you use it yourself.

Read-only, all the way down

Vibemetric never writes to ~/.claude or ~/.codex. The Claude run it starts can only Read, Grep and Glob. No shell, no edits, no web.

Secrets are stripped first

Before anything is read, the digest removes API keys, tokens, webhook URLs and database connection strings it recognizes.

Nothing lingers

The digest is deleted when scoring ends, and the scoring run is never saved to your Claude history, so it can’t skew your next score.

Questions

What does it cost?

The app is free. Scoring runs on your existing Claude plan through the claude command you already have installed, so it uses some of your plan like any other Claude Code session.

Which tools does it read?

Claude Code and Codex today, including Claude Code subagent transcripts. Cursor and Grok are next.

Why does scoring take five minutes?

The rubric only counts connected evidence: a rule saved in one session and used in a later one, or a failed check that was fixed and rerun. Finding those takes reading, not counting.

Is the score stable?

Scores come from an AI reviewer, so they can move a point or two between runs. Reasons always cite the sessions they come from, so you can check them.

Why does macOS warn me the first time?

Vibemetric isn’t notarized yet. Right-click the app, choose Open, then Open again. You only do this once.

Get your score.

Unzip, drag Vibemetric to Applications, then right-click and choose Open the first time. Needs Claude Code installed and signed in.

macOS 14+ · Apple Silicon and Intel · 1 MB download

Download for Mac