Research Questions研究问题
Notes toward a research agenda一份仍在形成的研究笔记
How can agents learn from experience to improve both their capabilities and their ability to learn and improve, moving toward recursive self-improvement (RSI)?
如何让智能体从经验中持续提升能力,并进一步提升自身学习与改进的能力,逐步实现递归式自我改进(RSI)?
The Goal: Improving the Ability to Improve目标:让改进能力本身也能改进
My long-term research goal is RSI. I am interested in the full learning loop: how an agent produces useful experience, obtains reliable feedback, and uses it to update model weights, memory, tools, and harnesses. The recursive part is a question to investigate: can an improved agent also improve how it generates experience, evaluates changes, and learns, making subsequent rounds of self-improvement more effective?
我的长期研究目标是实现 RSI。我关注完整的学习闭环:智能体如何产生有价值的经验、获得可靠反馈,并据此更新模型参数、记忆、工具与 harness。其中需要研究的“递归”在于:改进后的智能体,能否进一步改善产生经验、评价改动与学习的方式,使后续轮次的自我改进更有效?
A Key Question: Learning When Real Rewards Arrive Too Late一个关键问题:真实奖励来不及用于更新时,如何学习?
By extremely delayed rewards, I mean delays in real-world time. The outcome we care about may require months or years to become observable, while the model is already going through many training and update cycles. For recent examples whose labels depend on those future outcomes, ground truth is unavailable at training time. More computation alone cannot make the future outcome observable now.
这里的超级延迟奖励,指的是真实世界时间上的延迟。真正关心的结果可能需要数月或数年才能观察到,而模型在此期间已经经历了很多轮训练和更新。对于标签依赖这些未来结果的最新样本,训练当下无法获得 ground truth;仅仅增加计算,也无法让尚未发生的未来结果现在就变得可观测。
Suppose a label can only be determined after a fixed delay Δ. At training time t, only examples from t − Δ or earlier have mature labels. This creates a tension: labeled data describe the past, while the freshest data remain unlabeled. When the delay greatly exceeds the update cycle, true rewards from recent actions cannot close a timely training loop. Mature historical labels may still help with learning and calibration, subject to changes in the environment and policy.
假设一个标签必须等待固定时长 Δ 才能确定,那么在训练时刻 t,只有 t − Δ 及更早的样本具备成熟标签。这带来一个矛盾:有标签的数据描述过去,最贴近当前环境的数据却没有标签。当延迟远大于更新周期时,最新行为的真实奖励无法及时返回,当前更新必须面对这部分监督的缺失。历史上已经成熟的标签仍可能用于学习与校准,但必须考虑环境和策略已经发生的变化。
Mature Labels Still Need a Defensible Information Boundary历史标签成熟,不代表回测没有泄漏
For an LLM, a chronological split of downstream data does not establish what the model knew. A present-day model judging an old research idea may already have read papers reporting its eventual success. Future information can enter through pretraining, post-training, retrieval, or retrospective descriptions of the task. Hiding the label or rewriting the question cannot by itself establish a clean historical test.
对 LLM 而言,仅按时间划分下游数据,无法确定模型当时“知道什么”。让今天的模型评价几年前的研究想法,它可能已经读过报告该方向后来成功的论文。未来信息可以通过预训练、后训练、检索或事后整理的任务描述进入系统。只隐藏标签或改写题面,并不足以保证历史测试没有泄漏。
Using historical labels for supervised training is legitimate. The concern is whether the model learns shortcuts based on known outcomes and whether evaluation rewards the same shortcuts. Selecting only ideas known to have succeeded creates a separate selection bias. This leaves two difficulties: recent examples lack mature labels, while historical examples require evidence that the apparent foresight did not come from future knowledge.
用历史标签做监督训练本身是合理的。需要警惕的是,模型是否学到了识别已知结局的捷径,以及评测是否仍在奖励同一条捷径。只选择后来成功的想法,还会引入另一类选择偏差。于是困难有两面:最新样本缺少成熟标签,历史样本又需要证明表面的“远见”没有来自未来知识。
Within the goal of RSI, how can agents keep improving when true rewards are unavailable in time and historical supervision risks leaking future information, while providing credible evidence of that improvement? Learning effectively and establishing that learning occurred are related but distinct problems. Process reliability, predictions of future reward, and observed final outcomes need separate evaluation; a better proxy score alone does not establish a better eventual outcome.
面向 RSI,在真实奖励无法及时获得、历史监督又容易受到未来信息污染的条件下,智能体如何持续自我改进,并提供可信的改进证据?“怎样学习”与“怎样知道它真的学会了”是相互关联、又需要分别解决的问题。执行过程、未来奖励预测与最终结果应分别评价;代理分数变高,本身不足以说明最终结果会更好。
- What can support an update now? How can mature historical outcomes, intermediate observations, local experiments, and model predictions be used? What assumptions connect each signal to the long-term objective?当前更新可以依据什么?如何利用成熟的历史结果、中间观测、局部实验与模型预测?每种信号与长期目标之间,需要哪些关联假设?
- What makes evidence from historical data credible? How can we distinguish transferable judgment from recognition of known outcomes, and audit information available to the model as well as the task inputs?历史数据上的证据何时可信?如何区分可迁移的判断与对已知结局的识别,并同时检查模型已有知识和任务输入中的信息?
- How can learning stay relevant as the world changes? Can recent unlabeled experience help adapt what was learned from old labels, and how can the system detect when an old relationship no longer holds?世界变化后,学习怎样保持有效?最新的无标签经验能否帮助调整从旧标签中学到的规律?旧规律不再成立时,系统如何识别?
- How should unresolved outcomes affect exploration? An outcome that has not matured should remain unresolved. How can experience selection avoid favoring only quickly rewarded actions or treating pending outcomes as failures?未决结果应如何影响探索?尚未成熟的结果应保留为未决。经验筛选如何避免只偏向很快见效的行动,或把还没有结果的尝试当作失败?
- What can a reward correct when it finally arrives? Which records of information, model versions, actions, and alternatives are needed to revisit early judgments and credit assignments after many updates?奖励终于到来后,还能纠正什么?需要保留怎样的信息、模型版本、行动与备选记录,才能在多轮更新后重新评价早期判断和信用分配?
Research taste is one example. An idea's lasting value may become clear only long after the agent must choose which questions and experiments deserve investment. This setting also involves noise and disagreement about value. I see it as one setting for studying learning without timely ground truth within the broader goal of RSI.
Research taste(研究判断力)是其中一个例子。一个想法的长期价值可能很久以后才明确,但智能体必须提前判断哪些问题和实验值得投入。这个场景还包含噪声与对价值的不同判断。我希望把它放在 RSI 的总体目标下,用来思考缺少及时真实标签时的学习问题。
A Possible Route: Learning in Accelerated Simulations一条可能的路径:在可加速的模拟环境中学习
A sufficiently faithful simulator could compress months or years of waiting into much less computation time, making repeated trials and learning possible before real outcomes mature. Building a complete replica of the world is far beyond this proposal. I want to investigate a narrower requirement: make the simulator accurate on the questions, quantities, and mechanisms that determine the decisions we care about.
如果模拟环境足够可信,就可能把现实中数月或数年的等待压缩为更短的计算时间,在真实结果成熟前反复尝试和学习。完整复现真实世界远超这个设想的范围。我想研究一个更有限的要求:让模拟环境在我们关心的问题、数值,以及影响决策的机制上足够真实。
For example, a research agent may choose to scale up training, run an additional control experiment, or explore another direction. A useful simulator should capture how those choices affect eventual gains, resource use, and the evidence available for the next decision. Matching historical averages alone would not establish this: the simulator also needs to preserve meaningful differences between candidate strategies and the uncertainty around them.
例如,研究智能体可以选择扩大训练规模、补充对照实验,或探索另一个方向。有用的模拟器应当反映这些选择如何影响最终收益、资源消耗,以及下一步决策能获得的证据。仅仅复现历史均值还不够:候选策略之间有意义的差异,以及这些差异的不确定性,也需要得到保留。
- Responses to actions. Which relationships between actions and outcomes must remain faithful, and which details can be simplified without changing the decision?行动后的结果。哪些行动与结果之间的关系必须保持真实?哪些细节可以简化,而不改变应当作出的决策?
- Information timing. The training procedure can advance the simulated clock to obtain a terminal reward, but the agent must only see information available at each simulated decision time. Historical knowledge embedded in the agent or simulator also needs to be considered.信息出现的时序。训练程序可以快进模拟时间,获得终局奖励;智能体在每个决策时刻,只能看到那时可获得的信息。智能体或模拟器已有知识中包含的历史结局,也需要纳入检查。
- Validity after optimization. A simulator that predicts outcomes for familiar strategies may fail on new strategies found by the agent. How can we detect improvements that exploit simulation errors and test transfer beyond the training environment?优化之后仍然有效。模拟器可能准确预测熟悉策略的结果,却无法处理智能体新找到的策略。如何识别利用模拟误差获得的提升,并检验它们在训练环境之外能否迁移?
This connects to value equivalence, which defines model equivalence through Bellman updates for specified policies and value functions, and to MuZero, which learns predictions of reward, value, and policy for planning. These provide useful modeling ideas; they do not establish that a simulator can faithfully predict long-term research outcomes.
相关的建模思想包括 value equivalence(价值等价):针对指定的策略和价值函数,以 Bellman 更新是否一致来定义模型等价;以及 MuZero 对规划所需奖励、价值和策略的预测。这些思想可以借鉴,但不能据此认定科研的长期结果已经能够被准确模拟。
The central difficulty is calibration. Two simulators can fit the same short-term evidence while predicting opposite long-term effects. Without additional evidence or justified assumptions, more simulated trials cannot resolve that disagreement. Simulation rewards remain outcomes under a model's assumptions, not observed future ground truth. A possible learning loop would combine frequent simulated updates with checks from local experiments and, eventually, mature real-world outcomes.
核心困难在于校准。两个模拟器可能同样符合现有短期证据,却对长期效果作出相反预测。没有额外证据或有依据的假设,增加模拟次数也无法消除这种分歧。模拟奖励始终是模型假设下的结果,不能当作已经观测到的未来真实标签。一条可能的学习路径,是把频繁的模拟更新与局部实验检验结合起来,再由后来成熟的真实结果持续校准。
The Experience Loop经验闭环
I organize the path toward RSI around five related decisions: producing experience, building environments, evaluating behavior, assigning credit, and choosing updates. Extremely delayed rewards raise questions throughout this loop. These decisions can recur and interact as the agent learns.
围绕 RSI,我把这个闭环整理为五个相互关联的决策:产生经验、构造环境、评价行为、分配信用与选择更新。超级延迟奖励会贯穿这些环节;它们可以在学习过程中反复发生、相互影响。
- 01Define定义问题What experience?需要什么经验?
- 02Build构造环境How much reality?需要多真实?
- 03Judge评价行为Result and process结果与过程
- 04Attribute归因价值What deserves learning?什么值得学习?
- 05Update实施更新Where should it change?应该改进哪里?
What experience is worth producing?什么经验值得被产生?
Before implementing an environment, a task designer has to choose what capability or failure mode the task is intended to expose. Difficulty alone does not make a task useful. A stronger criterion may be whether success provides evidence about a capability that matters beyond a particular benchmark.
在实现环境之前,需要先明确任务希望考察什么能力,又希望暴露什么失败模式。难度本身并不足以说明任务有价值;一个更值得考察的标准,或许是任务表现能否反映某种不局限于特定基准的能力。
- Which task distribution represents the capability we actually care about?什么样的任务分布能够代表我们真正关心的能力?
- How should tasks evolve as the agent improves, instead of becoming a static test set?任务应如何随着智能体进步而变化,而不是退化成静态测试集?
- Can the agent help discover coverage gaps without becoming the sole author of its own examination?智能体能否帮助发现能力覆盖的空白,同时又不成为自己考试的唯一出题人?
What must an environment preserve from the real task?环境需要保留真实任务中的哪些部分?
For extremely delayed rewards, an environment could make future consequences available on an accelerated clock. The research question is which mechanisms must remain faithful for experience learned there to transfer. Under a limited budget, I want to study how accurately a simulator needs to represent action effects, information timing, costs, and uncertainty.
面对超级延迟奖励,环境可以尝试通过加速的模拟时间,让行动的后果更快可用。需要研究的是:哪些机制必须保持真实,才能让其中学到的经验迁移出去。在预算有限时,行动效果、信息时序、成本与不确定性分别需要模拟到什么精度?
- Which task-relevant structures and constraints must remain faithful for learning to transfer?为了使学习结果能够迁移,哪些与任务有关的结构和约束必须保持真实?
- How can we tell when a simplification has changed the capability being learned or evaluated?如何判断某种简化已经改变了原本希望学习或评价的能力?
- How should interaction, simulation, and verification budgets be allocated jointly?交互、模拟与验证预算应该如何联合分配?
What does the final outcome leave unverified?只看最终结果会遗漏什么?
A successful outcome can reflect robust behavior, luck, or evaluator exploitation; a delayed outcome may not yet be available for learning. I am interested in what observable process evidence can establish while the final reward is pending, and how later outcomes can test or correct the signals used for learning.
成功结果可能来自稳健行为、运气或对评价器的利用;延迟的结果则可能还无法用于学习。我关心在最终奖励未到时,可观察的过程证据究竟能说明什么,以及后来的真实结果如何检验或修正用于学习的信号。
- Actions, tool calls, state transitions, intermediate artifacts, and recovery behavior.动作、工具调用、状态变化、中间产物与异常恢复行为。
- Constraint satisfaction, robustness under perturbation, efficiency, and whether the agent asks for help at an appropriate time.约束满足、扰动下的稳健性、效率,以及智能体是否在合适的时机请求帮助。
- Evaluators that expose uncertainty instead of forcing every trajectory into a confident scalar reward.允许表达不确定性的评价器,而不是把每条轨迹强行压成一个自信的标量奖励。
What can be learned from a successful or failed trajectory?一条成功或失败的轨迹中,什么值得学习?
A long trajectory mixes decisive choices, harmless variation, recovery steps, and errors whose effects appear much later. When reward arrives after several policy updates, what was known and which model acted at each decision need to be preserved. I want to study how this history can support learning while distinguishing decision quality, execution failures, and outcome noise.
一条长轨迹混合了关键决策、无害差异、恢复步骤,以及很久以后才显现影响的错误。当奖励晚于多次策略更新才到来时,需要保留每次决策所依据的信息与模型版本。我想研究如何从这些历史中学习,同时区分决策质量、执行失败与结果噪声。
- Prioritize novelty, uncertainty, regret, failure coverage, and verifier confidence rather than reward alone.除奖励外,还应考虑新颖性、不确定性、遗憾值、失败覆盖与评价器置信度。
- Where controlled interventions are possible, compare alternative decisions from the same state and distinguish observed results from hypothetical ones.在允许受控干预时,从同一状态比较不同决策,并区分实际观测与设想中的结果。
- Separate policy failure from missing knowledge, tool failure, and evaluator ambiguity.区分策略错误、知识缺失、工具故障与评价歧义。
What should change to improve the next learning cycle?获得反馈之后,怎样改善下一轮学习?
Depending on what the evidence supports, an update may target model weights, memory, tools, the harness, or an evaluator. Beyond fixing a current failure, I want to investigate whether feedback can improve task generation, experience selection, and the update procedure itself, so that later rounds of learning become more effective.
根据证据支持的结论,适合更新的对象可能是模型参数、记忆、工具、harness 或评价器。除了修复当前失败,我更想研究反馈能否进一步改善任务生成、经验筛选与更新方法本身,使后续轮次的学习更有效。
- Weights: reusable policy or capability gaps.参数:可复用的策略或能力缺口。
- Memory: task-specific facts, precedents, and reusable experience.记忆:任务相关事实、先例与可复用经验。
- Tools and harness: repeated procedures, checks, and recovery paths.工具与 harness:重复流程、检查机制与恢复路径。
- Evaluator: missing constraints, ambiguity, and exploitable feedback.评价器:缺失约束、定义歧义与可被利用的反馈漏洞。
Three Working Hypotheses三个暂时的研究假设
Accelerated simulation may support transferable improvement.可加速的模拟,可能支持可迁移的改进。
Under a fixed budget, preserving the mechanisms that determine relevant decisions may be enough to learn useful strategies. This depends on how well those mechanisms are identified and whether the simulator stays reliable as the agent discovers new strategies.
在预算固定时,保留决定相关选择的关键机制,可能就足以学习有用的策略。这取决于能否可靠识别这些机制,以及智能体发现新策略后,模拟器是否仍然可信。
Recent unlabeled experience may help adapt learning from older outcomes.最新的无标签经验,可能帮助调整从旧结果中学到的规律。
Historical rewards and current observations may complement each other when their relationship remains informative. I want to test which assumptions make this possible, how proxy errors accumulate, and whether gains survive evaluation on outcomes that mature later.
当两者之间仍存在有效关联时,历史奖励与当前观测可能互相补充。我想检验这需要哪些假设、代理误差如何累积,以及改进能否经受后来成熟的真实结果的检验。
Some failures may be better addressed by changing the harness than the model.有些失败可能更适合通过修改 harness 处理,而非直接修改模型。
A system may improve more efficiently if it can locate the source of a failure and update the appropriate component. Whether this diagnosis can be made reliably remains open.
如果系统能够定位失败来源并更新相应组件,改进或许会更高效。不过,这种诊断能否可靠完成,仍有待验证。
Testing Without Future Information如何在不偷看未来的条件下检验?
A rolling evaluation should reconstruct the information available at each real-world cutoff: mature historical labels, recent observations, and unresolved outcomes. Save predictions and model versions at that time, then evaluate against outcomes once they mature. Separate an outcome's occurrence time from the time it became observable, and audit future knowledge already present in pretrained models when using historical data.
可以按真实时间滚动评估:在每个截止时刻,只提供当时已经成熟的历史标签、最新观测与仍未决的记录。保存当时的预测和模型版本,等结果成熟后再评价。同时区分结果发生时间与它真正可被观察到的时间;使用历史数据时,还需要检查预训练模型是否已经知道了未来信息。
Historical backtests need evidence about the model's training history, later adaptations, and retrieved sources, not just a date filter on the test set. Models trained with documented temporal boundaries, such as DatedGPT, offer one approach. Another is to freeze the system and record predictions before outcomes are known, as in ForecastBench. This reduces the risk of remembering resolved answers but retains the wait for real outcomes. These approaches address different parts of the tradeoff between evaluation speed, leakage control, and real-world relevance.
历史回测需要关于模型训练历史、后续适配与检索来源的证据,不能只给测试集加一个日期过滤器。像 DatedGPT 这样记录并控制训练数据时间边界的模型,是一种路径。另一种是先固定系统,在结果尚未知晓时记录预测,例如 ForecastBench 的做法;这降低了记忆已知答案的风险,却仍需等待真实结果。这些路径分别处理评估速度、泄漏控制与真实世界相关性之间的不同取舍。
With the same initial system, task distribution, and total budget, compare training on mature labels alone, learning from interim signals, combining historical labels with recent unlabeled experience, and adding accelerated simulation. Account for the cost of constructing, running, and checking the simulator. Later ground truth can evaluate updates that had to be made without it; it need not have been available as a training reward.
在相同初始系统、任务分布和总预算下,可以比较只使用成熟标签、依赖中间信号、结合历史标签和最新无标签经验,以及加入加速模拟的学习方式。模拟器的构建、运行与验证成本也应计入。后来的真实结果可以评价那些在缺少标签时已经完成的更新,并不要求它当时能够充当训练奖励。
For the simulation route, first check predicted returns and uncertainty for fixed candidate strategies, then test whether strategies improved inside the simulator transfer to held-out settings, mechanism changes, and prospective real tasks. Vary delay, noise, and distribution change separately. Artificially delaying known labels can test mechanisms; real-world long-term gains still need validation against outcomes that actually mature.
对于模拟这条路径,可以先检验固定候选策略的收益预测与不确定性,再检验经过模拟训练的新策略能否迁移到留出的场景、不同的机制设置,以及预先记录决策的真实任务。延迟、噪声和分布变化应分别控制。人为延迟已知标签可以检验机制;真实世界的长期收益仍需要实际成熟的结果来验证。
For RSI, I would also examine whether an update helps the agent learn more effectively in later cycles and transfer across tasks, accounting for additional computation, interaction, and verification costs.
围绕 RSI,还需要考察一次更新能否帮助智能体在后续轮次中更有效地学习、跨任务迁移,并计入额外的计算、交互与验证成本。
What I Have Not Figured Out我还没有想清楚的部分
- Who defines the next problem? Humans can provide purpose and boundaries; agents can expose blind spots. I do not yet know the right division of labor.下一个问题由谁定义?人类可以提供目的与边界,智能体可以暴露盲区;两者之间合适的分工尚不明确。
- What information makes learning possible before rewards arrive? Which assumptions connect observable signals to eventual outcomes, and how can the agent detect when that relationship changes?奖励未到时,学习依靠哪些信息?可观察信号与最终结果之间需要哪些假设?这种关系变化时,智能体又能否识别?
- How can the loop avoid self-confirmation? When an agent proposes tasks, builds evaluators, and learns from their feedback, external grounding and adversarial checks may be needed.闭环如何避免自我确认?当智能体同时提出任务、构建评价器并从反馈中学习时,可能仍然需要外部锚点与对抗性检查。
- What would demonstrate progress toward RSI? I want to distinguish a local task improvement, gains from additional resources, and an improvement that helps the system learn more effectively in later cycles.怎样验证正在走向 RSI?我希望区分单个任务上的局部改进、额外资源带来的收益,以及能够帮助系统在后续轮次中更有效学习的改进。
Further Reading on the Research Taste Example关于 research taste 例子的延伸阅读
- Kahneman & Klein (2009): Conditions for Intuitive ExpertiseKahneman 与 Klein(2009):专家直觉的形成条件
- AI Can Learn Scientific Taste (2026, v3)
These notes are still provisional.这些想法仍在形成。
I welcome discussions about sustained agent self-improvement, learning while true rewards remain unavailable, and cases where local gains fail to improve later learning. Concrete successes and failures can help identify useful mechanisms and assumptions that need revision.
欢迎讨论智能体如何持续自我改进、真实奖励不可得时如何学习,以及局部提升没能改善后续学习的案例。具体的成功与失败,有助于识别有用的机制与需要修正的假设。
