What Auto Research Lacks Is Not More Answers, but a Gate

Generation is no longer the most expensive part

In the past, any step in research—proposing a plausible hypothesis, finding relevant literature, writing a program, or completing an experiment—could take a long time.

Now AI can produce dozens of candidates in a single task, then generate mathematical derivations, code, experimental plans, and paper drafts. Give it more models, more agents, and more compute, and the number of candidates will continue to grow.

This makes it easy to think that the core problem of Auto Research is how to make AI generate more and better questions.

What I saw in the Minimax and MRI projects was the opposite: candidates were generated too quickly for validation to keep up; and scarcer still than validation was deciding which candidate deserved to consume validation resources.

Because even a candidate that eventually holds may be only:

Without an admission mechanism, increasing compute will not necessarily bring research closer to knowledge faster. It may only generate correct but trivial, complex but useless, or fundamentally indeterminate results faster.

What Auto Research lacks first is not more answers, but a gate that can say “no.”


This gate does not decide truth

I have defined the dual-track framework as Auto Research’s basic unit of governance: logical problems terminate in formalization, engineering problems terminate in reproducible results at the corresponding level, and problems that currently fit neither remain outside the boundary.

But the dual-track framework answers only this: by what mechanism should an already well-specified problem be decided?

It does not answer: is this problem worth consuming proof, experiment, data, equipment, and human attention?

These two matters cannot be conflated.

A proposition may be very important and ultimately be shown false; it can still be valuable research because it eliminates a genuine uncertainty. Another proposition may be rigorously proved in Lean and stably reproduced by a program, yet add nothing new.

So the meaning gate is not a third truth track within dual-track. It does not adjudicate “true” or “false”; it adjudicates something else:

Given the current goals, state of knowledge, and resource constraints, is this candidate worth sending into costly adjudication?

Passing the meaning gate does not mean a candidate is correct; holding up in the dual-track framework does not mean it is important.

candidate generation: what it might be
→ meaning gate: whether it is worth adjudicating
→ dual-track: whether it holds within its own type
→ pass through the gate again: whether this result is worth further investment

The gate governs resources; dual-track governs propositions. They are adjacent and cannot substitute for one another.

One current boundary also needs to be stated: the meaning gate below is a governance design I have abstracted from Minimax’s experience of stopping and MRI’s record of negative results. It is not a process that has already been fully deployed and shown effective in the MRI project. The two projects supplied the problems and local practices; the gate itself still needs to be implemented, revised, and tested in more research.


Minimax: when should complexity stop increasing?

Minimax did not begin by calling this a “meaning gate.” It simply gave me my first clear encounter with a situation where a small signal remained, yet continuing might no longer be meaningful.

The project tried six versions of small models in succession. It later organized the 64 hexagrams and 384 lines of the I Ching into data, studying how much textual variation traditional structures could explain.

At first there was a seemingly exciting correlation coefficient: r≈0.28. It would have been easy to misstate that as “explaining 28%.” After adding cross-validation that held out entire hexagrams and random baselines, the explanatory rate the project could stably retain narrowed to about 14%, and it came mainly from positional structure. Measured with minimum description length, the gain was only about 0.06–0.11 bit per line.

These figures did not prove that the I Ching contains no further structure, nor did they establish an absolute upper bound. They described only the data, features, models, and evaluation protocol at that time.

The question was: should we keep piling on more complex models?

If the goal is merely to make a number keep rising, more features, interaction terms, and parameters can always be added. If the goal is to reduce genuine uncertainty, then multiple models and measurements had already roughly converged, the added complexity brought no commensurate engineering advantage, and the remaining question required different material or a different observational perspective.

So the project stopped.

It did not stop because the result was negative. More precisely, the project obtained a bounded result and also found that the marginal information from further investment along the same path was already low.

The meaning gate here did not address whether “14% counts as success.” It addressed whether the next unit of resources could still materially change our understanding of the problem.


MRI: only after correctness does the real problem begin

The MRI project pushed this tension further.

After a cellular sheaf sampling candidate was connected to a finite-dimensional HARDI model, the related logical propositions were formalized and the numerical implementation could be reproduced. Yet the resulting sheaf bandwidth was, in this model, simply the familiar 1/condition number.

It may be correct.

It may also fail to yield a new MRI metric or actionable guidance.

Another attempt hoped to use this metric to choose regularization for underdetermined reconstruction. The experiment ultimately exposed a structural problem: the metric could describe conditioning over the effective range, but it could not see the null space and therefore could not resolve the real non-identifiability. Continuing to optimize the same metric would not suddenly give it the ability to observe the null space.

Literature checks found that some other candidates already had mature precedents, or that after formalization and numerical scans they collapsed into condition-number phenomena. They were not all false propositions. Some held entirely within their stated scope, but did not make a sufficient new contribution.

These results made me realize that Auto Research cannot consist only of a “propose—prove—reproduce” pipeline.

If a system upgrades every formally correct, reproducible engineering result into a “research outcome,” it will eventually produce a large number of indisputable but irrelevant fragments of knowledge. The stronger the validation, the more these fragments may even resemble genuine breakthroughs.

The meaning gate must appear before costly adjudication, and it must appear again after a result holds.


What must a candidate answer before entering validation?

Meaning cannot be compressed into a universally applicable score. But a candidate can be required to submit an admission card that answers at least six questions.

1. Which real unknown does it address?

A candidate must point to a gap, contradiction, unexplained phenomenon, or default assumption in the knowledge map. Merely “applying A to B” is not a problem; it must explain what exactly remains unresolved in B.

2. What does it add relative to existing work?

Changing terminology, notation, and narrative does not add knowledge. The system should actively search for the closest existing results and explain whether the addition is a condition, boundary, prediction, construction, or decision-making capability.

3. What remains after the new packaging is removed?

Perform an ablation: remove the purportedly new structure and return to an ordinary baseline. If the result barely changes, the added structure may be only decoration.

4. What new observable consequence does it produce?

A candidate need not immediately produce a product or clinical result, but it should at least change something checkable: a prediction, a boundary, a counterexample, a construction, or a trade-off between two directions.

5. What result could kill it?

If every output can be interpreted as support, the candidate has not yet entered research. Stopping conditions should be written as clearly as possible before seeing the result.

6. What is the cheapest discriminating experiment?

Do not first write an entire paper or formalize a cathedral. First seek the smallest experiment, counterexample, or theorem that most sharply distinguishes “worth continuing” from “should stop.”

These six questions are not intended to guarantee a positive result. They are intended to ensure that, whether the result is positive or negative, the investment can reduce some explicit uncertainty.


The gate must not become a dictatorship of taste

Once we begin discussing what research is worth doing, danger follows.

Meaning depends on goals, context of use, and costs. A mathematical interface that offers nothing new for MRI engineering may be valuable for teaching formalization methods; an experiment that is too costly today may become worthwhile after equipment changes. So-called importance cannot be fully derived from a theorem.

If the gate rewards only short-term, quantifiable returns, it will systematically exclude the most unfamiliar, hardest-to-describe, and potentially most paradigm-changing questions. If it obeys only the established tastes of senior experts, it may instead entrench a field’s current blind spots.

So the gate is neither an immutable ranking nor the intuitive verdict of one person. At a minimum, it needs to:

This means that the meaning gate is itself an object of governance.

It will reject some things wrongly, and let some trivial results through. Its value is not in always choosing correctly, but in making the reasons, failures, and revisions behind resource allocation traceable.


Rejection should also become a research artifact

Traditional papers usually show only the results that passed every gate. If Auto Research keeps records in the same way, it will face a serious problem: the system will not see why it once stopped, and will therefore regenerate the same candidates in the next round.

Therefore, being rejected by the gate should not mean being deleted.

Every rejection should leave at least:

the candidate and its source
the goals and resource conditions at the time
the closest existing work
which checks it passed
which gate stopped it
the evidence for stopping
what change would be sufficient to reopen it

Minimax’s value lies not only in the final approximately 14% or 0.06–0.11 bit, but also in recording why model complexity was no longer increased. MRI’s negative results did not only close off several candidates; they also made null space, condition number, and semantic mapping boundaries that future candidates must confront early.

For an automated system, these records are also a memory against repetition. They ensure that a failure does not have to be paid for again after switching to another agent or another wording.

Negative results are not inferior versions of positive results. As long as they eliminate a previously genuine possibility and their boundaries and evidence are sufficiently clear, they reduce the unknown.


Compute should amplify depth after filtering, not quantity before filtering

I was once constrained by compute, so I naturally imagined that more resources would make it possible to produce research results more quickly and in greater volume.

That may be only half true.

More compute can indeed search more fields, propose more candidates, and complete more proofs and experiments. But as the number of candidates grows, source checking, semantic translation, independent reproduction, real-world feedback, and value judgment do not automatically grow at the same rate. Without a gate, compute may first amplify the queue awaiting validation and common-mode errors.

Truly scalable Auto Research should not send every idea into its most expensive process. It should first use the cheapest evidence to eliminate candidates that are obviously repetitive, unfalsifiable, lacking a new contribution, or inconsistent with current goals, then concentrate resources on the few paths that still change our understanding.

This gate will not tell us what we should ultimately study. It can only make every “continue” state its reasons, and every “stop” leave evidence.

Research remains full of contingency. Candidates that pass the gate may all still fail, and the most valuable discovery may come from a direction that initially seems strange.

The gate is not meant to eliminate that contingency.

It is meant to ensure that random exploration does not mean consuming resources without memory, and that high-speed generation does not mean high-speed declaration of results.

When answers can increase without limit, the truly scarce capability may not be continuing to answer, but knowing which question is qualified to consume the next unit of real-world resources.