A shared vocabulary for records that people and machines can inspect. Amaze amaze amaze.
In Andy Weir's Project Hail Mary, Ryland Grace and the alien engineer he calls Rocky must build a way to communicate across radically different bodies and environments. Their shared vocabulary grows through attention, correction, and work they need to do together. The fiction supplies a memorable image for a real design problem: a fluent exchange is useful only if the participants can establish what they mean.
An AI conversation presents a different kind of distance. A person brings a life, a purpose, and an understanding of the situation. A model generates an answer through learned computational processes. Its fluency can help the exchange, but fluency alone does not establish that it has preserved the person's meaning.
Representational AI needs a bridge that remains inspectable. Natural language is essential to that bridge. Structured records can help hold distinctions that a conversation leaves implicit and give reasoning systems explicit objects to work with. Mirad, a constructed taxonomic language, is one candidate worth investigating for part of that work. The larger aim is a semantic foundation on which different models can operate, with a vocabulary chosen for how well it helps people express and inspect their experience.
What a conversation can lose
“My team” can refer to four collaborators, a department, or everyone who helped over a season. A person may move between those meanings quite naturally. A system preparing an account for an employer needs to know which one applies.
“I might try this” is different from “I have decided to do this.” A rehearsal of an opponent's argument is different from agreement with it. A joke about perfectionism is not a diagnosis or a stable personality trait. If those distinctions disappear during summarization, the resulting memory can be fluent and wrong.
Persistent memory and retrieval can carry information between conversations. They can also carry a mistaken interpretation forward. The design problem is therefore more specific than remembering more: store the appropriate account, with its source and status, and make correction effective wherever that account is used.
Natural language can name, distinguish, and correct these things. We do it constantly. Structured identifiers and explicit fields can make some of the work more dependable across software components. In the proposed substrate, those distinctions would also become things the system reasons over: whether two accounts concern the same event, which claim a correction supersedes, or whether an intention has become a decision. Conversational flexibility would help us work with that structure without having to speak in database fields.
Why Mirad interests me
Mirad descends from Noubar Agopoff's Unilingua project and has been developed and documented by Jamie Shoemaker. Its reference grammar describes a deliberately organized vocabulary rather than a lexicon assembled solely through historical inheritance. Relationships among words are meant to carry relationships among concepts.
That design makes it interesting for a bounded representation task. Could a small, systematic vocabulary help people and software keep distinctions stable across retellings? Could it make the relation between an action, an observation, an interpretation, and an intention easier to inspect? Those are questions a prototype could test.
Calling such words coordinates is an analogy. A taxonomic address is not automatically a coordinate chart on a mathematical manifold. To make that stronger claim, we would need to define the space, the mapping, and the properties that follow. The language's organization does not provide those results by itself.
Nor is a constructed vocabulary culturally neutral merely because its rules are regular. Categories express choices about what matters and how things belong together. An account of family, responsibility, conflict, or achievement may resist the distinctions an ontology offers. The person needs a way to say that the vocabulary is wrong for the experience.
My interest in Mirad is practical and provisional. Shoemaker's work offers material to examine and a starting point for collaboration. It deserves testing against alternatives, including ordinary language labels with stable identifiers, existing ontologies, and a small task-specific schema. The test should reward usefulness, not allegiance to the candidate.
A bounded vocabulary for a bounded job
The first vocabulary need not represent all human thought. Its job could be as limited as keeping a Dote's action, participants, observations, evidence, interpretation, questions, and permissions distinct. Additional concepts should enter because participants need them, with a documented process for revision. Participants need a way to challenge a translation, propose a category, and understand how revisions affect older records. Versioned mappings can preserve differences between communities without forcing everyone into one classification.
A compact vocabulary would not make every experience compact. Some accounts need a story, an image, a silence, or competing descriptions. The structured layer should point to those materials and preserve their context. It should never imply that the categories exhaust the person.
Consider Maya's robot repair again. The system needs an identifier for the particular session, the participants she chooses to name, a link to the test log, and the distinction between “we found a loose connector” and “I wonder whether a checklist would help.” It does not need to decide that Maya is permanently a leader, a mechanic, or a conflict-averse person.
If the Rep later prepares an application paragraph, each material claim should be traceable to the approved account. A vocabulary might help align the records across systems. It cannot authorize sharing, prove that the account is true, or make a model's interpretation into Maya's own judgment.
Gameshow as a place to test the bridge
Participants should be able to speak, write, or show what happened in their own terms. Gameshow would help assemble a proposed record and present the consequential distinctions in language they understand. Learning a constructed language should not be a condition of participation.
The interface might ask, “Is this something you observed, or your explanation of what happened?” It might show who is named in a story and ask which parts may be shared. It might leave a field unresolved because the participant does not yet know. These are ordinary acts of clarification made visible in the record.
Human confirmation also has costs. A system that asks people to approve every tiny inference can produce fatigue and habitual agreement. A pilot must find which decisions deserve attention, which defaults are understandable, and how a person can recover from an approval they regret. A checked box is not sufficient evidence of informed control.
One useful experiment would compare a plain-language schema with a Mirad-assisted one on the same limited tasks. Can participants correct the record? Can they find the source of a statement? Do translations preserve stance and reference? Does the vocabulary reduce errors enough to justify its complexity? Negative results would be valuable: they would tell us to use a simpler bridge. The original account should remain available under the participant’s control so that a translation can always be checked against it.
Implication without false authority
Once approved records can be assembled and retrieved reliably, a Rep may help explore their implications. Several Dotes might suggest that a person enjoys a particular kind of work, avoids a recurring difficulty, or learns well with a certain collaborator. These should begin as proposals the person can examine, not conclusions about who they are.
I use implication models to name the longer ambition: intelligence developed to reason over the structure of lived experience. In such a system, events, claims, intentions, and their relationships would be primary computational objects. Learning could concern how experiences unfold and which relationships survive a retelling; reasoning could compose those objects into a proposed account or course of action. Neural learning and symbolic reasoning offer approaches to investigate together. The choice of architecture should follow what these tasks require, with authored memory and human correction shaping its development.
Suppose Maya corrects an account that credited her with proposing a test her teammate suggested. The system should be able to locate which later interpretations relied on that claim and reconsider them. Her contribution may still show care, persistence, or a willingness to learn, but the reasoning now has a different starting point. An early prototype can expose those dependencies in records and rules. A more ambitious model could learn and reason directly over them, with its proposed interpretations returned to Maya for reflection. That is the progression from a useful memory service toward the AI of semantic representation.
An implication should remain linked to the accounts that prompted it and be marked as an interpretation. Rejecting it should not erase the original accounts. Approving it for reflection should not authorize its disclosure to a school, employer, or insurer. These distinctions matter more than which vocabulary a developer finds elegant.
The architecture should also allow uncertainty about other people. A person's sincere account is evidence of their experience, not unrestricted authority to define everyone in the story. Witnesses can disagree. Records involving several people need rules for attribution, access, and contest, not a single unquestionable version.
The bridge must remain open
The image from Project Hail Mary stays with me because shared work makes the vocabulary necessary, and patient correction makes it useful. Our systems need that same room for correction. A person should be able to say, “That is not what I meant,” and change what the machine carries forward.
The goal is an account that can travel without quietly becoming someone else's account of us, and intelligence that can reason from it without losing the relationships that give it meaning. A shared vocabulary may help build that foundation. The authority to question and revise what it carries is what makes the bridge worth crossing.
Sources
Andy Weir, Project Hail Mary (Ballantine Books, 2021). The fictional communication problem supplies an analogy, not evidence for the proposed representation system.
Noubar Agopoff, Unilingua — Langue universelle auxiliaire (Paris, 1966), is the original publication of the language now called Mirad. Agopoff’s project was to construct an auxiliary language on principles drawn from mathematics and chemistry rather than from natural-language inheritance.
Mirad Grammar, including “Introduction,” the public reference grammar documenting the language developed by Jamie Shoemaker from Agopoff’s project. Consulted September 2026: https://en.wikibooks.org/wiki/Mirad_Grammar
On neurosymbolic AI, see Artur S. d’Avila Garcez and Luis C. Lamb, “Neurosymbolic AI: The 3rd Wave,” Artificial Intelligence Review 56 (2023), pages 12387–12406, and Henry Kautz, “The Third AI Summer” (AAAI Robert S. Engelmore Memorial Lecture, 2020).
Michael Robbins, “Beyond Inference: Implication Models and the Future of Human-Centric AI,” discussion draft, April 2025 (W3C public-humancentricai archive). Introduces implication models, the DOTES schema, and the use of Mirad as a symbolic representation layer, and credits Jamie Shoemaker’s translation and modernization of the language from Agopoff’s original French. https://lists.w3.org/Archives/Public/public-humancentricai/2025Apr/att-0005/Beyond_Inference_Discussion_Draft_April2025.pdf