Sjá · · 7 min read

Sjá 0.8: the probe learns to say why

A citation probe that says 'not cited' sends you off to do one of two opposite jobs at random. Sjá 0.8 adds three judgements — coverage, fidelity and stance — so the probe stops being a scoreboard and starts being a diagnosis.

Sjá 0.8: the probe learns to say why

Since 0.6, Sjá has been able to ask the question that matters most in GEO and that nothing else in the field will answer for you: when a user asks ChatGPT, Claude, Perplexity or Gemini the question this page answers — does the engine cite you?

You type the question. Sjá puts it to the four engines with your own keys, concurrently, and judges each answer: cited or not, at what position, and who was cited instead. Every probe is a row, failures included. Schedule it and you get a trend, and an alert the day an engine stops citing you and names who took your place.

That's been useful. It has also had a hole in the middle of it, and 0.8.0 is about the hole.

The dead end

A probe comes back not cited. Now what?

You genuinely don't know. There are two completely different situations hiding behind that one word:

  1. The page doesn't actually answer the question. It's about the topic, it ranks for the keyword, but a user asking exactly this would not find their answer in it. The engine is right to look elsewhere. This is a content problem, and the fix is writing.

  2. The page answers the question fine, and the engine chose someone else anyway. Bigger site, more inbound links, a name the model has seen more often. This is an authority problem, and no amount of rewriting the page will fix it — the work is elsewhere, and it's slow.

Same verdict. Opposite jobs. And a probe that says only no sends you off to do one of them at random. We watched people rewrite pages that were already good, and leave pages alone that were genuinely thin, because the tool couldn't tell them which was which.

Coverage

0.8.0 asks a second question, of a different kind of model. Alongside the probe, Sjá hands the page's own main text and the user's question to Jev — TypeSafe's judgment model, which answers typed questions about text rather than generating it — and asks: how completely does this page answer this question?

The answer comes back as a score, and the card now reads:

Not cited — ChatGPT, Claude, Perplexity
Coverage: 0.91 · the page answers the question
→ This is an authority problem. The page is fine; the engines chose other
  sources. See who: nhs.uk, mayoclinic.org, healthline.com

or:

Not cited — all four
Coverage: 0.34 · the page does not answer the question
→ This is a content problem. Missing: the question asks about eligibility;
  the page covers how to apply and never says who qualifies.

Two different cards, two different jobs, named. That's the whole feature, and it's the one that stops the ten hours of rewriting a page that didn't need it.

Fidelity

Then there's the other case nobody had been checking: cited, and wrong.

An engine cites you — good — and in the same breath says your product costs $49 when it costs $39, or that you support Windows when you don't, or quotes a return policy you changed two years ago. From the citation count this is a win. From a customer's point of view it's your name attached to a false statement, in the most authoritative-sounding voice on the internet.

Fidelity checks each thing a citing engine says about the site against the page:

Cited — Gemini, position 2
Fidelity: misrepresented
Gemini says: "Elyra Sjá is available for macOS and Windows"
The page says: macOS on Apple silicon; there is no Windows build

The claim, the page's own words, side by side. And if this is a scheduled probe, the first time a citation goes from faithful to misrepresented raises an alert — because that's the one you want to know about the day it happens, not the month after a customer asks why you lied.

Stance

The third judgement is the simplest and, it turns out, the one people had wanted longest. An engine can mention you in four quite different tones:

  • recommends — "the best option here is…"

  • lists — one of several, no preference

  • references — cites you as a source without opinion

  • advises against — "…though you may want to avoid…"

Since 0.7 Sjá has recorded mentioned, no link as its own category, because a mention is neither a citation nor an absence. But a mention that says "avoid" and a mention that says "the best" were the same row. Stance separates them, for every engine, cited or not — so the mentioned count finally says whether the mention helps.

How it decides, and how sure it is

Jev is not a chat model. You don't prompt it for prose; you give it a state — here, the page's text and the engine's answer — and typed questions: a score for coverage, a yes/no for each claim's fidelity, a choice for stance. It returns calibrated probabilities. That matters for two reasons.

First, it's what lets an answer become a card rather than a paragraph you have to read and interpret. authority_problem: 0.88 maps to a sentence and a next step. A model's essay about your page doesn't.

Second, it's how Sjá knows when to say nothing. A stance is only stated from 0.5 confidence. If Jev can't tell whether Claude is recommending you or merely listing you, the card shows the answer and no stance — because a confidently wrong advises against would send you into a panic over an answer that was neutral. The probe already said what to read is the trend, not one answer; the same humility applies to the judgement of the answer.

One more honest line, which the interface shows: Jev reads English best. A Norwegian page will be judged, and probably judged reasonably, but the scores are calibrated on English and the interface says so rather than letting you find out.

What leaves the machine

Sjá has always been strict about this, and adding a judgement model is exactly the moment to be stricter.

Everything above is opt-in. There's a field for a TypeSafe key in Settings; until you put one in, nothing changes. Probes run exactly as they did in 0.7 — your own engine keys, your own account, nothing else.

With a key, the boundary is stated where you add it, not on a page you'd have to go looking for:

With a TypeSafe key, the probed page's main text and the engines' answers
leave this machine for TypeSafe, to be judged. Without one, probes work
exactly as before.

That's the sentence. The page's text is public — it's the page — but "public on your website" and "sent to a third party by a tool on your laptop" are different facts, and you should get to decide the second one deliberately.

And the judgements are shown and stored, never acted on. Sjá doesn't rewrite your page because coverage was low, doesn't stop probing an engine because it advised against you, doesn't do anything at all with Jev's answer except put it on the card and in the history. You read it. You decide.

Why this matters

GEO has a measurement problem that SEO never had. In search, you could see the results page; the ranking was the feedback. With answer engines you get a paragraph, and the paragraph is a black box: cited or not, and no reason. The citation probe in 0.6 opened the box far enough to see the verdict. 0.8 opens it far enough to see why — not the engine's actual reasoning, which nobody has, but the three questions whose answers tell you what to do next: does the page deserve the citation, is the citation telling the truth, and does the mention help or hurt.

Three judgements, each with a probability, each opt-in, each shown beside the evidence it was made from. The probe stops being a scoreboard and starts being a diagnosis.

Get it

Sjá 0.8.0 is at elyracode.com/sja — macOS on Apple silicon, notarized, updated in place from the app. Already installed? Check for updates. The changelog has the full list, and the security page has the boundary in one paragraph for whoever needs to approve the key.

Run a probe you've already seen come back not cited. Add a TypeSafe key. Run it again, and read which of the two jobs it turned out to be.