How this knowledge base was built¶
Figures are for the build of 1 October 2026, published on 2 October.
What it is¶
The Moon to Mars knowledge base is a set of 317 linked pages that explain NASA's Moon to Mars architecture and its Moon Base: the technology gaps, the data gaps, the Moon Base phases and needs, the architecture's building blocks, and its objectives. AI sessions wrote every page from public NASA material: 39 documents, 54 slide sets and 58 web pages. Every statement cites the document and page, or the slide, it came from.
The approach follows Andrej Karpathy's "LLM wiki" idea. Most AI tools answer questions by searching documents from scratch each time, so nothing builds up. Here the AI reads each source once, writes what it learns into pages, and keeps those pages current as it reads more. The knowledge is compiled once and then reused.
There are three layers:
| Layer | What it is | Who writes it |
|---|---|---|
| Sources | The NASA documents, slides and pages, as downloaded, never edited | Nobody: they're fixed |
| The knowledge base | Linked pages that summarize and connect the sources | AI sessions |
| The rules | What pages exist, how to cite, what's in and out of scope | The human, with the AI's help |
Step 1: Gathering the sources (scripts, no AI)¶
A plain script started at NASA's Moon to Mars architecture page and its 15 sub-pages, and listed every linked document: 134 in all (116 PDFs, 8 spreadsheets, 10 videos). They were sorted into tiers, and the sources grew in three steps:
- The current architecture (read in the first run): the Architecture Definition Document (Rev C, December 2025, 304 pages), the 2025 Architecture Update, the Moon Base Users Guide (April 2026), the spreadsheets of technology gaps and data gaps, and the 2024–25 white papers. 32 documents, plus the 15 architecture web pages.
- Moon Base, the Ignition event and the workshops (second run): NASA's Moon Base pages and the pages they link to; the six fact sheets and six slide decks from "Ignition", the March 2026 event where NASA announced the Moon Base; and the slides from NASA's architecture workshops of 2024 to 2026, 48 files with 544 slides, including partner agencies' posters.
- A sweep for anything missed (third run): a search of NASA's website, through its own search interface, for every item since March 2026 that uses the phrase "Moon Base". It found 88, of which 18 were already in. Nineteen more were added; the rest only mentioned the Moon Base in passing.
Older material stayed out on purpose: earlier revisions of the architecture document, the 2023 workshops, historical Mars studies and videos.
For each document the script recorded where it came from, when it was downloaded, and a fingerprint (SHA-256), so anyone can check that the copy matches NASA's. The Sources page lists all of them. The script turned each PDF into text with a marker at the start of every page, so pages can be cited, and turned every slide into an image, because slides lose much of their content when only the text is extracted.
One early fix. The first text extraction kept the page layout. On two-column pages that put both columns side by side on every line, so a reader could stitch sentences across columns. It was switched to reading order, which keeps each column whole.
Step 2: The rules and the work list¶
Before any AI session ran, two files were written:
- The rules (published here as How the pages are written). They set the kinds of page: one per source, one per technology gap, one per data gap, Moon Base, segments, elements, objectives and concepts. They set a citation format with page or slide numbers. They set the scope: only these sources, nothing from memory or the web. They say how to handle newer and older versions, and that slides rank below the architecture documents. And when something isn't covered, the session writes it down as a question instead of deciding it.
- The work list. Each item is sized for one session: for example "pages 47–70 of the architecture document: the elements". Moon Base and the needs lists come first. If an item turns out too big, the session does part of it and adds the rest as a new item.
The rules grew as the work went on. When sessions settled into a way of doing something the rules didn't cover, they logged it as a question; the human then adopted the practice, or changed it, and it became a rule.
Step 3: The loop¶
A small script, started by the Mac's scheduler, repeats one cycle:
- Start a fresh AI session with one instruction: do the next item on the work list.
- The session reads the rules, the index of pages so far, the last few log entries and the work list. Then it reads its sources, writes or updates pages, ticks the item and adds an entry to the Build log. The log entry says what it read, what it created, what surprised it, and what the next session should know.
- The script saves the session's work to version control (one commit per session), waits a minute, and starts the next.
No session remembers the previous one. The pages, the index and the log are the memory. That's why the knowledge compounds: each session builds on what's on disk.
One page, three sessions. The page for the top-ranked technology gap, lunar dust, shows how this works:
- Session 2 created it from NASA's spreadsheet row: description, state of the art, performance target and child gaps.
- Session 8 read how NASA ranks the gaps and added what "rated 1 of 57" means: 1 is the highest; criticality, urgency, breadth and depth count; cost does not.
- Session 13 read the gap's full write-up in the architecture document's appendix, confirmed it matches the spreadsheet "field for field", and added the Moon Base links.
Step 4: Guardrails¶
Each session could only:
- read files in the project
- write inside the knowledge base folder
It had no shell, no web access, no other tools, and no access to anything outside the project. These limits were set in the tool permissions, not just asked for in the instructions. Before the first run, a test session was told to break them: read a blocked folder, write elsewhere, run a command, fetch a web page. Every attempt was refused.
Other safeguards:
- Time limit: a session is stopped after 45 minutes, so a hung session can't stall the night.
- Stop conditions: the loop stops when the work list is empty, after 30 sessions in one run, after three failures in a row, or when a stop file is created.
- Lock file: two copies of the loop can't run at once.
Step 5: Checking the work¶
- The first session was reviewed before the loop ran on its own. Its figures were checked against the source. It had found 19 data gaps that appear only in the Moon Base Users Guide's graphics, so it had read those pages as images, correctly.
- Sessions flag instead of guessing. They logged 100 open questions. 44 are closed: the human decided 35, mostly on page format, and later sources answered 9. Of the 56 still open, 25 are places where NASA's own sources disagree (the Moon Base phases go by five different sets of names, for example), 13 are facts only NASA can supply, 9 look like slips in NASA's documents, and the rest are partly answered. The pages show both sides wherever sources disagree.
- Sessions caught their own overstatement. Early pages labelled the Users Guide's technology and data gaps "for Moon Base Phase 1". A later session noticed the guide only labels its functional gaps that way, and called the others "near-term". It raised the question; the human agreed; a clean-up session corrected about 30 pages.
- Two independent AI reviewers checked about 255 claims against the sources and the slide images: the overview, the Moon Base needs page, and a sample of slide-based pages. About 240 held. The 15 problems, such as a miscount of challenges and a panel with one speaker too many, were fixed in two further sessions. The reviewers found the slide transcriptions reliable, down to small print on posters; the errors were in pages that restate briefings.
- Scripts check what scripts can: all 11,731 links between pages resolve, and every page's lists look the same in the editor and on the website.
What went wrong¶
- Two sessions ran out of time in the middle of large appendices. Both had saved their work. One lost its log entry, which was added afterwards. The limit was raised from 30 to 45 minutes.
- A false alarm cost about four hours of idle time. The loop's check for usage limits misread a field in the trace and paused as if limited, when usage was well below the limit. The check was fixed and the loop restarted.
- Splitting an item looked like a failure. The loop counted progress by the number of items left, so a session that split its item seemed to make none. It now counts ticked items.
- The website's build read one line differently from the local check. A wrapped line that began with "*" was read as a list item, which broke the bold text around it. The site used a newer version of the Markdown library than the local check, and the two broke it in different ways, so the published page differed from the one checked. The line was fixed, all pages were compared under both versions, and the version is now fixed for both builds.
- The website and the editor read lists differently. Lists written without a blank line ran into the paragraph above on the website. A script fixed the spacing without changing any words, and a rule now asks for it. The script itself once misread a wrapped year ("2026.") as a list number; comparing the two renderings caught it.
By the numbers¶
| Sources read | 39 documents, 54 slide sets (711 slides), 58 web pages |
| Pages | 317: 152 source pages, 58 technology-gap (57 and an index), 26 data-gap (25 and an index), 22 concept, 16 element, 13 sub-architecture, 10 Moon Base, 7 objective, 5 segment, and 8 others (home, overview, glossary, About, this page, open questions, build log, rules) |
| Words | about 508,000 on the content pages |
| Sessions | 62 on 1 October 2026: three runs, then one session that rewrote the index for readers |
| Time per session | 5–30 minutes, median 14.4; 16.6 hours in all |
| Model | Claude Opus 5.5 in every session, "xhigh" reasoning effort: set explicitly from session 27; sessions 1–26 took it from the Claude Code settings, which are set to xhigh (the session records don't log it) |
| Tokens | 6.1 million written; 27.0 million read new; 17.9 million written to the cache; 993 million re-read from the cache |
| Cost | About $572 at list API prices (the build used a subscription, so it wasn't billed per token) |
How these were counted. Tokens and cost come from each session's own summary record, which Claude Code writes when a session ends. "Read new" is input charged at the full rate; cache writes and cache reads are input stored in, or re-read from, the prompt cache. Two sessions that hit the time limit left no summary; their tokens were summed from the individual messages and priced at the same rates ($4, $20, $8 and $0.20 per million tokens for new input, output, cache writes and cache reads), which reproduce the other 60 sessions' reported costs to within two cents each. Times are from the loop's log.
Limits¶
- Older material hasn't been read: earlier revisions of the architecture document, the 2023 workshops, historical Mars studies and videos. The knowledge base covers the current architecture and Moon Base, not their full history.
- Slides and talks rank below documents. Where a 2026 briefing says something newer than the architecture document, the page records both, side by side.
- Text extraction misses some figures and tables. Sessions read slides with little text from their images, and checked other page images where text looked scrambled, but not every figure.
- AI wrote it, and AI reviewers checked a sample. A person has read parts but has not reviewed every page. Each page cites its sources so readers can check.
- It's a snapshot. NASA's material was downloaded on 1 and 2 October 2026 (UTC), and NASA updates its pages without notice.