The most useful part of the House ASI archive is not its earliest confidence. It is the later review that asks what the implementation actually did.
That review identifies several gaps between a component’s name and its behavior. Publishing those corrections makes the research easier to assess and gives the next implementation a more honest starting point.
A project name is not a result
“House ASI” describes an ambition. The initial implementation was a multi-role orchestration prototype with selected deterministic solvers.
The revision therefore removes claims of demonstrated superintelligence or emergent capability. Those claims require evidence beyond the existence of an architecture.
The useful work remains: organizing generation, evaluation, selection, and records so the proposal can be tested.
A cache is not neural test-time learning
The reviewed “Titans Memory” implementation used ordinary stored records and deterministic routing. The critique reports no learned neural memory module or gradient-based test-time update.
That matters because the Titans paper describes a neural long-term memory approach. Inspiration from its terminology or hierarchy is not an implementation of its mechanism.
The current component should be described according to what it stores and retrieves. A future neural implementation would need its own validation.
A hash is not a semantic embedding
The archive describes SHA-derived numbers labeled as embeddings, while decoding recovered the original text from metadata.
A hash can identify content and support integrity checks. It is not designed to place similar meanings near one another.
The working contribution was a packet-transfer structure. Calling that “latent communication” or meaningful semantic compression obscured the real behavior.
Multiple roles need a comparative test
The reviewers found no empirical evidence that the initial four-role pipeline improved over a simpler call on the same tasks.
The accepted response proposed a comparison with a single call and a multiple-candidate baseline. That is the right next experiment, not a result that can be marked complete in advance.
Framework comparisons need the same discipline. A feature table assembled from aspirations and selected implementation details does not establish superiority to another system.
Operational errors belong in the model
The review also identified errors flowing as candidate text and weaknesses in evaluation. A system must distinguish valid outputs, invalid outputs, provider failures, and unavailable evidence.
These are engineering issues with concrete fixes. They should not be hidden beneath a confidence score.
Use current memory research to sharpen the terminology
The original Titans paper proposes a neural memory mechanism that learns from context. The August 2026 Lychee Memory V2 preprint instead studies a pipeline for organizing and retrieving persistent memories. These are different mechanisms even though both can improve use of past information.
For every component, name the stored artifact and the operation that changes it. A database row update, a prompt revision, an embedding index rebuild, and a gradient update should appear as different events in the evidence record.
Then ask what changed in the delivered answer or action because of that event. If removing the component leaves all tested outputs unchanged, investigate whether the test is insensitive or the component is inactive.
Make corrections testable claims
Convert the review into a small matrix: original claim, observed implementation, corrected description, proposed repair, and a test that could establish the repair. Keep “proposed” visible until the test has actually run.
For semantic retrieval, use similar meanings with different wording and unrelated statements sharing keywords. For candidate diversity, inspect outputs and solution strategies rather than counting role names. For verification, supply plausible wrong answers that the evaluator must reject.
Repeat the comparison with the component disabled. Report the effect on quality, latency, and failure handling rather than assuming that a repaired implementation must improve everything.
A newer paper can suggest a better experiment; it cannot retroactively establish that the archive implemented that paper. The strongest update is therefore a more precise vocabulary and a reproducible test plan, while preserving the historical record of what the review actually found.
What should survive the correction
Keep the component decomposition, explicit evidence records, domain routines that pass appropriate checks, and the willingness to inspect failure.
Remove unsupported score trajectories, universal ceilings, invented semantic mechanisms, and claims that a larger graph proves greater intelligence.
A correction is not a substitute for the next experiment. It is what makes the next experiment worth trusting.
Keep a good idea close.
Follow Signal Tower


