How to orchestrate five AI systems as a unified cognitive architecture, with you as the sovereign consciousness.
Just as you know when to use your hands versus your eyes, substrate literacy is knowing which AI substrate fits which task.
"This awareness only emerges at the Hexad level. Single-AI users never develop it because they only have one tool."
Match tasks to the AI whose neural analog specializes in that cognitive function.
When passing work between AIs, context must be preserved. The transfer protocol ensures nothing is lost.
The human is not a "user"—they are the cognitive sovereign. AI amplifies. Human decides.
The Command Center shows scores, prices and context windows for five models. This section defines every one of them — where it came from, when it was read, and which of them are opinions.
Scores and prices are sourced. Roles and analogies are editorial. Anything we could not source is shown as “not sourced” — never as a low score, and never as an empty bar. A filled bar is the visual grammar of measurement, so only sourced values get one.
An earlier version of this page rated each model on reasoning, creativity, code, safety, speed and cost from 0–100. Those numbers had no rubric and no source — they were estimates presented in the costume of measurements. They have been removed rather than refreshed, because a confidently-rewritten estimate is harder to distrust than an obviously stale one.
Read off a primary source. Carries the source URL and the date it was read, or it is null.
Computed from Tier 1 in code. Nobody types these in, so they cannot drift from their inputs.
The 7:2:1 weighting is Artificial Analysis's convention, published here so the figure is reproducible.
If a vendor publishes no cache rate, blended is “not sourced” — the input price is never substituted for it. Standing input in for cache would price that model as though caching saved it nothing, inflating its cost against rivals with nothing in the output to show it happened.
Where a cache figure is derived from a stated policy rather than published as a rate — OpenAI publishes a 90% cached-input discount, not a price — it is marked *, and everything computed from it carries the same mark.
Our opinions, and useful ones. They are labelled editorial wherever they appear and are rendered as text or chips — never as a bar.
“Claude is the Prefrontal Cortex” is a teaching metaphor, not a finding. Treat the routing advice the same way.
For running models as an ensemble rather than picking one. Architecture, base weights, corpus and training stack — categorical facts, not scores.
A node that scores well but shares a substrate with another node adds less than a weaker node that fails differently. Most vendors publish none of this, so most of these fields read Undisclosed — which is what the vendor discloses, not a low rating.
The Artificial Analysis Intelligence Index is revised. A score from v4.1 set beside one from v4.1.1 is a wrong number that looks right. Every score on this site was read from a single page at a single version, and each stores the exact model variant it came from.
That last part matters more than it sounds: the index lists one model at several reasoning-effort settings (Claude Opus 5 appears at 63, 61 and 59). Aggregators quoting “the” score for a model are usually quoting a different row from the one you would use. That is why we name the variant.
Selection rule: highest available variant per model. The variant names on this site are not uniform — three read (max) and two read (high) — because the settings offered differ by model. Checked 2026-08-13: (max) exists for Claude Opus 5, GPT-5.6 Sol and DeepSeek V4 Pro; it is not offered for Gemini 3.7 Flash (high/medium/low) or Grok 4.6 (high only), whose own index pages are titled “Gemini 3.7 Flash (high)” and “Grok 4.6 (high)”. So every score here is that model's top setting.
We state this because without it you cannot tell whether the gap between a 63 and a 56 is capability or configuration — and comparing a low setting of one model against a high setting of another is one of the ways this page has been wrong before.
Every row carries the date it was read. A build check flags any row older than 60 days, and flags pricing notes whose stated change-date has passed. The original problem here was not carelessness — it was that nothing could detect the drift. An undated row cannot go stale; it can only be wrong.
Model data last read 2026-08-13 · Intelligence Index v4.1.1 (released 2026-08-06).
Every figure comes from the vendor's own documentation or from Artificial Analysis directly. Third-party price aggregators were used only to cross-check, never as the cited source — several of them contradict each other and one another's index versions.
One benchmark index is not model quality. It does not measure how a model handles your domain, your prompt style, your latency budget or your tolerance for a particular failure mode. It is a comparable, dated, checkable starting point — which is a different and smaller thing than a verdict. The routing advice on this site remains a judgement call, and it is labelled as one.
What are you trying to accomplish? Be specific. The clearer the mission, the better the routing.
Use the Task Router in the Command Center, or develop your substrate literacy to route intuitively.
Work with the AI. When it hits limits, transfer context to another. Keep the human anchor in the loop.
The human anchor integrates all outputs, applies judgment, and makes the final call. AI advises. Human decides.
Your gut says no, even if all AIs agree. Trust it.
AIs process ethics. You embody them. You decide.
You know things AIs don't. Your lived experience counts.
Human relationships require human judgment.