Questions

Does the judge know which model wrote what?

No. Passages go to Claude Opus labelled A, B, C… in an order derived from the trial, with no model names anywhere in what it reads. It does see your prompt and your goal, because it has to score compliance against them.

Isn't Claude judging Claude a conflict?

It would be if it could tell. It can't — and you can see for yourself: if you include a Claude model in the trial, its passage is judged under a letter like all the others. When the verdict picks it, read the reasoning and decide whether you agree.

What are the numbers?

A fixed ruler run over every passage: words per sentence, the p10 and p90 sentence lengths, how much the length varies, fragments, sentences that open with a dependent clause, conjunction density, verb-to-adjective balance, interior thought, paragraph size, dialogue share and tagging. They are counts, not opinions. The judge is given them as ground truth so it quotes real figures instead of guessing.

Wait, what is a p10 and a p90?

Two numbers from the table that describe the range of sentence lengths in a passage. Line every sentence up from shortest to longest. The p10 is the length one-tenth of the way along: 10% of the sentences are that short or shorter, so it tells you how clipped the passage gets. The p90 is the length nine-tenths of the way along: 10% are that long or longer, so it tells you how far the long sentences stretch. A passage with a p10 of 3 and a p90 of 30 swings between fragments and long, winding lines; one with a p10 of 9 and a p90 of 15 keeps an even pace. Neither is good or bad on its own — it depends what your prompt asked for — which is why the judge is given them and told to quote them when they bear on the verdict.

How is the novel cost worked out?

Each call reports the tokens it used and the gateway's price for them. We divide that by the words the passage produced and multiply by 50,000. It assumes a novel is written as many prompts about this size, each re-sending a prompt about this size — which is how most people actually work. It is an estimate, not a quote.

How is the price set?

A flat fee based on your prompt's length ($1.00 up to 25,000 words; $2.00 up to 100,000 words; $3.00 up to 200,000 words), which covers the Claude Opus assessment and other costs. Then each model you picked at its provider's list price for the run — every token of your prompt, the full thinking ceiling for the level you chose, and a 1,000-word passage — with no markup. Card processing is passed through at cost. You see it broken down by model and can revise before paying.

Why do the model lines look high for a short scene?

Because they are the ceiling. A model allowed to think 16,000 tokens may think 2,000, and the line assumes it uses all of it. That's the price either way; if you'd like it lower, pick a lower thinking level.

What if a model fails?

You're not charged for it. Your card is authorised for the quoted total and captured for the quote recomputed without any model that failed to return text. If every model fails, the authorisation is released and the trial is marked failed.

What settings do the models run with?

The thinking level you chose (Off, Standard, Deep or Max — a reasoning-token ceiling), temperature 1.0, and passages capped at about 1,000 words. No system prompt of ours — your prompt is the whole thing.

Is there a limit on prompt length?

200,000 words, which is a long novel. Within that, a model whose context window can't take your prompt is left off the quote with a note saying so. The judge reads the whole prompt too; that's what the fee tiers are for.

How long do you keep my prompt?

30 days, then it's deleted along with the passages and the assessment. The transaction record stays for accounting. See the privacy policy.

Can I get a model added?

Probably. Email support@aiproselab.com with the model's name.