SENN Content NetworkIdeas worth building on.Explore the network
Research Notes

An AI research budget you can audit

Separate setup, training, inference, retrieval, and review costs before accepting an efficiency claim.

Revised and condensed from the recovered Studio7 drafts. Historical project claims are distinguished from independently verified facts; the original experiments were not rerun for this edition.

A blank job ledger and document folders on a workshop bench beside stacked lumber.
AI-generated editorial image · Signal Tower

A research budget should tell you what experiment you can afford. It should not turn a collection of optimistic ratios into a promise of frontier capability.

The recovered economics chapter proposed a remarkably cheap training run, then compared it with estimates for much larger systems. It also multiplied gains attributed to tokenization, retrieval, data filtering, and model architecture. Those calculations were proposals, not evidence that the complete system achieved equivalent quality.

Define the unit of useful work

Begin with a workload and an acceptance rule. For example: answer a fixed set of technical questions with correct citations, or resolve a collection of software issues that pass independent checks.

Measure cost per accepted result. A system that generates inexpensive text but requires extensive correction can be more expensive to operate than its token bill suggests.

For training, state the model, dataset, number of runs, evaluation schedule, and intended capability. Comparing a small experimental training job with a frontier development program is not meaningful unless the delivered capabilities and accounting boundaries are comparable.

Use a complete ledger

Record compute hours, hardware allocation, electricity or rental charges, storage, data preparation, failed runs, evaluation, and human review. Separate money already spent from the resources consumed by the next experiment.

An owned machine may reduce immediate cash spending. Its time is still finite. Include a separate marginal-cost view if that is the decision you need, and label it clearly.

Keep provider rates and hardware assumptions dated. Do not publish the archive’s old price table as current purchasing guidance.

Efficiencies interact

Fewer tokens do not automatically imply proportionally less total training cost. A changed vocabulary can alter embedding size, sequence lengths, data processing, learning behavior, and the metric being compared.

Retrieval can reduce some work while adding indexing, storage, queries, and maintenance. Sparse activation changes compute differently from total parameter storage. Faster generation can be offset by more candidates or verification passes.

Measure each proposed change alone, then together. A combined result is an experimental question. Multiplying headline numbers from different papers does not answer it.

Budget for a decision, not a victory

Fund the smallest experiment that can reject the idea. Specify the baseline, acceptance threshold, failure threshold, and maximum expenditure before starting.

Keep enough budget for repetition and independent evaluation. A single successful run with no remaining resources to reproduce it is a fragile research outcome.

The next stage should depend on evidence from the previous one. If an encoding scheme saves space but reduces downstream accuracy, the budget must allow the project to stop or change direction.

Separate construction savings from operating savings

The August 2026 revision of Faster Superword Tokenization addresses tokenizer-training efficiency. Its construction-time improvements should be booked against data preparation, not multiplied directly into every later training and inference expense.

Likewise, the August 2026 Lychee Memory V2 preprint reports more efficient memory construction in its experimental setup. A deployment still needs to account for retrieval, storage, correction, and the effect on task success.

Create separate ledger rows for one-time setup, periodic rebuilding, and each query or run. When a technique moves work between these categories, report the workload volume at which the tradeoff becomes worthwhile.

Price one complete experimental decision

A useful budget unit is a comparison that can support a go-or-stop decision. Include the baseline run, candidate runs, failed attempts, independent evaluation, and the time needed to inspect disagreements. Reserve funds for reproducing the selected result.

For example, estimate cost per accepted answer as total experiment expense divided by answers meeting the predeclared criterion. Also report the success rate and sample size. A low ratio from a tiny, selected sample is not a reliable forecast.

Keep cash expenditure separate from allocated hardware time and human review hours. Present uncertain inputs as ranges. If electricity price, provider rate, or utilization changes, the reader should be able to recompute the result without reconstructing a spreadsheet’s hidden assumptions.

Do not use current-looking currency figures unless their source and date have been checked. This update supplies an accounting method rather than a fresh price quote or a new claim about the archive’s training costs.

Publish the assumptions with the result

A useful report includes a compact ledger, measured quality, elapsed time, and sensitivity to the uncertain inputs. Readers can then substitute their own electricity rate, hardware constraints, or review cost.

The affordable-research ambition remains worthwhile. It becomes credible when another builder can follow the accounting and understand precisely what the experiment bought.

03

Keep a good idea close.
Follow Signal Tower

Keep reading

A few more good questions.

A place in your reading list

Good ideas, at your pace.

Follow Signal Tower in your favorite feed reader. No inbox required.

Follow the journal