Conversation
Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
…and efficiency Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
|
(torch) PS D:\work_caiyingwei\projects\Algorithms-Examples> cd '.\PPO with vanilla RNN\src' Starting training for 48 updates (500000 timesteps)Update 10/48 | Timesteps: 102400 | Avg Reward: 0.0000 | Policy Loss: 0.1415 | Value Loss: 0.9932 | Entropy: 0.6925 ============================================================ |
Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
Fixed in d6c4d84. The visualization crash occurred because no episodes were completed during training (each trial is 5500 time steps, but the default rollout length is 512). The Note: With the default settings, episodes take longer to complete than the rollout length. For longer training runs or adjusted parameters, episodes will complete and the full performance metrics will be displayed. |
|
@copilot 我发现代码还有一些细节问题:ax3.boxplot中似乎应该使用labels而不是tick_labels;训练过程中Avg Reward一直是0.0,即使lick_cost很高、reward-value很低也显示0.0,应该是环境代码中重置导致的;训练使用的默认参数导致"no full episodes were completed",需要调整为合适的参数; |
…ility Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
Fixed in 4628d2c. Issues addressed:
|
…ollection Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
|
@copilot 从agent的表现来看,虽然学会了在刺激窗口lick,但在非刺激窗口也总是lick,这是一种"作弊"的方法,请思考并尝试修改代码使得agent学会”到了刺激窗口再lick“的行为,此外,不再设置probe trial以减小任务难度,避免对agent的学习产生混淆干扰。 |
…ove action sampling Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
…or learning Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
Fixed in 49b40f6. Changes made:
Current status: The reward structure now correctly incentivizes the optimal behavior (licking only during stimulus gives ~60 reward vs ~0 for always-licking). However, the biologically-constrained continuous-time RNN is still learning slowly. The agent achieves ~1-3 reward, better than never-licking (-7.5) but not yet optimal. Further tuning of the RNN architecture (tau, spectral radius) or using a simpler model may be needed to achieve the desired "lick only during stimulus" behavior. |
|
@copilot 目前的任务对于"PPO with vanilla RNN”可能太难了,改成先从简单的任务开始(可以参考pdf文章中的某个任务,例如Standard Task任务),先正确复现算法确保"PPO with vanilla RNN”可以使agent学习到合理的策略,相关的文档说明也需要对应修改。 |
Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
Fixed in 4cfceca. Implemented a simpler Standard Task (标准刺激-反应任务): Task description:
Changes:
The optimal strategy is straightforward: respond when stimulus is present, don't respond otherwise. This makes it ideal for validating the PPO + vanilla RNN framework. Training shows entropy decreasing (0.69 → 0.43) indicating learning is happening. Further hyperparameter tuning may be needed to reach optimal performance (~15 reward per episode). |
…tion Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
…-choice Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>
已对比参考代码仓库 (https://github.com/aldobattista/neural-circuit-economic-choice),发现并修复了以下关键差异 (516e86a, 61041aa): RNN 动力学修复:
PPO参数对齐:
发现的问题: 建议:
|
…rove comments Co-authored-by: 1re2turn1 <149225935+1re2turn1@users.noreply.github.com>



Implements PPO-based training for biologically plausible vanilla RNN, starting with a simple Standard Task (stimulus-response) to validate the algorithm before moving to more complex timing tasks. Implementation aligned with reference code from neural-circuit-economic-choice.
Environment (
environment.py)Model (
model.py)h = α * ReLU(W_rec @ h + W_in @ x + b + noise) + (1-α) * h√(2σ²/α)PPO (
ppo.py)Training & Visualization
train.py: CLI-configurable training with checkpointing and task type selectionvisualize.py: Training curves, behavior analysis (hit/miss/FA/CR rates for Standard Task), network activity heatmapsKnown Issues
The Standard Task provides a simpler learning objective to validate the PPO + vanilla RNN framework works correctly, with clear optimal strategy that the agent can learn.
Original prompt
✨ Let Copilot coding agent set things up for you — coding agent works faster and does higher quality work when set up for your repo.