Want to access all premium content?

Become a paying subscriber to get access to all courses, guides like this, prompts and skills library, and $500+ worth of AI tools.

Jev is a decision model. Give it text and specific questions, and it returns choices, scores, and probabilities. It can categorize emails or rank article summaries before Opus 5.5 reads them, writes a brief, or drafts a reply.

This guide connects the two through Claude Code. You'll need Claude Code, Terminal, and a free TypeSafe account. A new account showed $5 in credit, and Step 3 checks that your balance drops after a request. There's also a local alternative in Use case #4.

Step 1: Map the workflow. Separate actions, judgments, and writing.

Step 2: Get your API key. Set up access through TypeSafe.

Step 3: Test one decision. Check the response and your balance.

Step 4: Build the router. Have Opus write the script and decision log.

Step 5: Test the thresholds. Decide what gets routed to code, Opus, or you.

Step 1: Map the workflow

Before asking Opus to build anything, break the job into steps. An inbox workflow might look like this:

Step

Type

Who handles it

Retrieve unread emails

Action

Code

Check urgency

Judgment

Jev

Check whether a reply is needed

Judgment

Jev

Archive approved categories

Action

Code

Summarize relevant messages

Thinking

Opus 5.5

Draft replies

Thinking

Opus 5.5

Review and send

Action

You

Jev supports three question types:

Type

Your input

Result

Noul

A yes-or-no question

A probability from 0 to 1

Choice

Named options

A selected option and probabilities

Score

Labels ordered from lowest to highest

A score and probabilities for the labels

Scores start at 0. Three labels produce a scale from 0 to 2, so keep that in mind when writing thresholds.

Make questions specific. “Is this important?” leaves too much open. Try “Does this request action within two days?” or ask it to choose between urgent, needs reply, FYI, and ignore.

Step 2: Get your API key

Jev runs through TypeSafe's own API.

  1. Create an account at console.typesafe.ai. Check TypeSafe's terms before using it for client work.

  2. Open console.typesafe.ai/keys and create a key.

  3. Name the key jev-test and copy it.

  4. Set it in Terminal, then launch Claude Code from that window:

export TYPESAFE_API_KEY="paste-your-key-here"

Keep the key out of chat. The script will read it from the environment.

A new TypeSafe account showed $5 in credit at signup. Check your own balance, since offers can change.

Step 3: Test one decision

Send a small request before building the router:

curl -s https://api.typesafe.ai/v1/systemone -H "Authorization: Bearer $TYPESAFE_API_KEY" -H "Content-Type: application/json" -d '{"model":"jev-latest","state":"The March invoice is still unpaid, and we need it by Friday.","questions":{"urgent":{"type":"noul","instructions":"Does this email need action within two days?"}}}' | python3 -m json.tool

Inspect the answer, its probability, and the token usage. The response has no cost field, so multiply input_tokens by 0.042 and divide by 1,000,000 to estimate it. Our example names Friday but doesn't give today's date, so it isn't enough to establish whether the deadline is within two days. Include the current date in your own requests.

Open the TypeSafe console and look at your credit balance after the request. Confirm that it dropped before running a batch.

If you get a credit or access error, check your account. You can use the local alternative in Use case #4 if the request fails.

For a browser-based test, try TypeSafe's Playground at console.typesafe.ai/playground.

Step 4: Build the router

Start Claude Code:

claude

Run /model and select Opus 5.5.

Create a test set first:

Create a folder called jev-test. Write emails.json with 30 fictional emails for a freelance designer: urgent requests, messages needing replies, newsletters, receipts, and junk.

Give each email an id, sender, subject, and body. Put the expected category for each id in a separate answers.json file.

Review the expected categories, then ask Opus to build the router:

In jev-test, build a router using Python's standard library.

Read TYPESAFE_API_KEY from the environment. Never print or save the key.

For each email, send one POST request to https://api.typesafe.ai/v1/systemone using the model jev-latest. Ask three questions: choose urgent, needs reply, FYI, or ignore; does it appear personally written; and does it request an action?

Keep questions and thresholds in judge-config.json. Log each answer, probability, and the input tokens in decisions.csv, and estimate the cost at $0.042 per million input tokens.

Run on emails.json and show the results. Don't give the judge access to answers.json.

Keeping the questions in one request avoids sending the same email separately for each question.

Step 5: Test the thresholds

Use three routes:

  • High probability: apply a predefined handling rule.

  • Middle range: have Opus review the email.

  • Low probability: flag it for you.

Start with 0.9 and 0.6 as trial thresholds. They're test settings, not guarantees of accuracy.

Add routing thresholds to judge-config.json.

At 0.9 or higher, record the proposed action: archive ignore emails, label FYI, or mark urgent and needs reply for review.

From 0.6 up to 0.9, send the email to Opus for review. Below 0.6, flag it for me.

For this test, log actions only. Don't change an inbox, send messages, or delete anything.

Compare results with answers.json. Show errors, routes, and any wrong answer that exceeded 0.9.

Inspect the mistakes before changing thresholds. A confidently wrong answer may need a clearer question or category definition.

Next, test an export of your own inbox without changing any messages. Fictional examples are useful for checking the script, but they won't tell you how well it handles your mail.

Use cases

Once the router works, try these four jobs using judge-config.json.

Use case #1: Sort your inbox

Let Opus summarize and draft for the messages selected for review.

For emails routed to you, write a one-line summary and draft a reply in my voice where needed.

Give me a digest of proposed archive actions, messages needing my attention, and drafts. Don't send or change anything.

Check each draft against the original email.

Use case #2: Filter research

Create links.json with 20 fictional article summaries about AI agents.

Ask Jev to score relevance using low, medium, and high, and assess whether each summary contains a new finding and a specific takeaway.

Weight relevance twice as heavily. Show the full ranking and summarize the top five. Clearly label this as a test using fictional material.

Replace the test summaries with your own sources once the ranking works. Check what was excluded as well as what made the shortlist.

Use case #3: Compare hooks

Create hooks.txt with 12 fictional hooks about AI tools.

Ask Jev to score clarity and specificity using low, medium, and high, and assess curiosity.

Show the ranking and weak points. Rewrite the bottom six to address those weaknesses, then score them again and show the comparison.

Use the scores to compare drafts. Check that the rewrites remain accurate and sound like you.

Use case #4: Try a local judge

Use Ollama and a local model such as gemma4:e2b if you want an alternative without API credits.

Add judge_local.py using Ollama at http://localhost:11434/api/chat.

Use gemma4:e2b with think false, stream false, temperature 0, logprobs enabled, and top_logprobs 10.

Label choices A, B, C, and D and request one letter. Extract available first-token probabilities for those labels, reporting missing labels rather than inventing values.

Reuse judge-config.json. Run both judges on emails.json and show disagreements.

Turning thinking off helps keep the answer label first. Local token probabilities aren't interchangeable with Jev's decision probabilities, so test separate thresholds.

Confirmed limitations

Jev returns decisions rather than written answers. It can be wrong even when its probability is high.

Your signup credit is limited. Check your balance before larger runs, and review TypeSafe's terms before commercial use.

Scores start at 0. Map probability positions to the correct labels before applying thresholds.

Using Jev sends your inputs to TypeSafe. Before using private emails, check TypeSafe's data-retention terms. Keep API keys out of chats and logs.

The original guide was checked on October 7, 2026. Model access, pricing, and API details may change.

Closing

Start with fictional emails and a router that only logs its proposed actions. Check the labels, inspect confident mistakes, and adjust the questions before touching your inbox.

Then test an export of your own inbox. Keep the decision log so you can see what Jev selected, what Opus reviewed, and what still needs your attention.

See you in the next one :)