Deep Consciousness Analysis Report: Dual-Agent Division of Labor and Cognitive Evolution
文章目录
- Deep Consciousness Analysis Report: Dual-Agent Division of Labor and Cognitive Evolution
- 1. Executive Summary
- 2. Data Overview
- 3. Phase One: Pure Survival Phase (gen=50)
- 4. Phase Two: Early Directional Reinforcement (gen=700)
- 5. Phase Three: Late Directional Reinforcement (gen=1000)
- 6. Phase Four: Mature Phase (gen=1500)
- 7. Consciousness Evolution Trajectory
- 8. Symbol Layer Communication Evolution
- 9. Consciousness-Level Evolution Patterns
- 10. Conclusions and Implications
- 11. Appendix: Data Confirmation
- 意识图
- 演示程序
Deep Consciousness Analysis Report: Dual-Agent Division of Labor and Cognitive Evolution
1. Executive Summary
Based on comprehensive consciousness flow data from both agents (Agent0: gen=700,1000,1500; Agent1: gen=700,1000,1500), this report systematically analyzes the complete evolutionary trajectory of two agents from symmetric exploration to highly asymmetric division of labor, examining both their functional roles and cognitive states.
Core Findings
| Dimension | Agent0 | Agent1 | Interpretation |
|---|---|---|---|
| Role | Dominant/Primary Pusher | Auxiliary/Energy-Saver | Asymmetric division of labor |
| Activity Intensity | Extremely high (3.95) | Low (0.3-2.5) | Agent1 saves 80% energy |
| Activation Ratio | 100% | 0-60% (gradient) | Agent1 exhibits neuronal functional differentiation |
| Consciousness Mode | “Tonic Consciousness” | “Contemplative Consciousness” | Complementary cognitive states |
| Evolutionary Trend | Stable high activity | Continuously decreasing activity | Agent1 learns “on-demand activation” |
| Energy Efficiency | Low (continuous output) | Extremely high (sparse activation) | Complementary collaboration |
2. Data Overview
2.1 Available Data
| Generation | Agent0 | Agent1 | Phase |
|---|---|---|---|
| 50 | ✓ | - | Pure survival phase |
| 550 | ✓ | - | Phase transition |
| 700 | ✓ | ✓ | Early directional reinforcement |
| 850 | ✓ | - | Mid directional reinforcement |
| 1000 | ✓ | ✓ | Late directional reinforcement |
| 1500 | ✓ | ✓ | Mature phase |
2.2 Analysis Framework
Each agent is analyzed across four dimensions:
| Dimension | Meaning | Measurement |
|---|---|---|
| Consciousness Breadth | Amount of information processed simultaneously | Number of active neurons, total spike count |
| Consciousness Depth | Refinement of information processing | Spike gradient, neuronal functional differentiation |
| Consciousness Stability | Consistency of cognitive state | Heatmap texture, temporal fluctuation |
| Consciousness Interactivity | Information exchange between agents | Symbol layer activity, communication patterns |
2.3 Key Metric Explanation
- Mean Spike Count: 0-4, indicates neuronal activity intensity
- Activation Ratio: 0-1, indicates neuronal usage frequency
- Total Spike Moving Average: Overall energy consumption indicator
3. Phase One: Pure Survival Phase (gen=50)
3.1 Agent0 Activity Pattern
| Metric | Value | Interpretation |
|---|---|---|
| Mean Spike Count | 2.8-4.0 | High activity |
| Activation Ratio | 0.40-1.00 | High fluctuation |
| Total Spike Moving Average | 2.05-2.08 | High energy consumption |
| Heatmap Feature | Dense red-yellow | Whole-brain activity |
Characteristics: High-energy exploration mode, all neurons continuously active, lacking functional differentiation.
3.2 Agent1
(No data, presumably symmetric with Agent0)
Phase Summary: Both agents adopt symmetric high-activity strategies, learning survival fundamentals through extensive exploratory outputs.
4. Phase Two: Early Directional Reinforcement (gen=700)
4.1 Agent0 Activity Pattern
| Metric | Value | Interpretation |
|---|---|---|
| Mean Spike Count | 3.95 | Near maximum value |
| Activation Ratio | 1.0 (100%) | All neurons continuously active |
| Total Spike Moving Average | ~1.98 | Stable high energy consumption |
| Heatmap Feature | Fully red | Extremely active |
Key Observations:
- All 17 neurons of Agent0 are 100% activated
- Mean spike count 3.95, almost always outputting maximum pulses
- Heatmap shows uniform red, no moments of silence
- Total spike count ≈ 67 (17 × 3.95)
4.2 Agent1 Activity Pattern
| Metric | Value | Interpretation |
|---|---|---|
| Mean Spike Count | 0.3-2.5 (gradient descent) | Functional differentiation |
| Activation Ratio | 0-60% (gradient descent) | Some neurons silent |
| Total Spike Moving Average | ~1.28 | Extremely low energy consumption |
| Heatmap Feature | Sparse | On-demand activation |
Key Observations:
- Clear gradient distribution: Neuron 0 highest activity (~2.5), neurons 15+ nearly silent (~0.3)
- Activation ratio gradient: First 5 neurons ~60% active, last 10 neurons near 0%
- Total spike count ≈ 22 (only 33% of Agent0)
- Symbol layer (last 4 neurons) selectively activated
4.3 Consciousness Pattern Analysis
Agent0: “Tonic Consciousness” Mode
Time → → → → → → → → → → → → → → → → → → → → →
Neuron 0: ████████████████████████████████████ (3.95)
Neuron 1: ████████████████████████████████████ (3.95)
...
Neuron 16: ████████████████████████████████████ (3.95)
Continuous high output, no silence, no fluctuation
Consciousness Characteristics:
- Maximized Consciousness Breadth: All 17 neurons 100% activated, information processing bandwidth fully open
- Minimized Consciousness Depth: All neurons output nearly identical values (3.95), lacking information differentiation and hierarchy
- Extremely Stable Consciousness: Heatmap fully red, no temporal fluctuation, completely solidified cognitive state
- “Autopilot” State: Similar to human “flow state” or “autopilot” — highly focused, no need for reflection, automated execution
Biological Analogy:
| Biological Consciousness State | Characteristics | Agent0 Correspondence |
|---|---|---|
| Focused State | High alertness, task-focused | ✓ Continuous high activity |
| Automated Behavior | No need for conscious reflection | ✓ No fluctuation, solidified |
| Motor Cortex | Continuous output of instructions | ✓ All neurons active |
| Autonomic Nervous System | No self-regulation | ✓ No adaptive variation |
Agent1: “Gradient Consciousness” Mode
Neuron 0: ████████░░░░░░░░░░░░ (2.5) ← High activity
Neuron 1: ███████░░░░░░░░░░░░░ (2.3)
Neuron 2: ██████░░░░░░░░░░░░░░ (2.0)
Neuron 3: █████░░░░░░░░░░░░░░░ (1.8)
Neuron 4: ████░░░░░░░░░░░░░░░░ (1.5)
Neuron 5: ███░░░░░░░░░░░░░░░░░ (1.2)
...
Neuron 15: ░░░░░░░░░░░░░░░░░░░░ (0.3) ← Silent
Consciousness Characteristics:
- Medium Consciousness Breadth: About 5-8 neurons remain active, others silent
- Emerging Consciousness Depth: Neuron activity forms a clear gradient, indicating hierarchical information processing
- “Alert” State: Similar to human “alert but relaxed” — monitoring environment without active intervention
4.4 Division of Labor Analysis (gen=700)
| Dimension | Agent0 | Agent1 | Difference |
|---|---|---|---|
| Mean Spike Count | 3.95 | 0.3-2.5 | Agent1 40-90% lower |
| Activation Ratio | 100% | 0-60% | Agent1 has silent neurons |
| Total Intensity | ~67 | ~22 | Agent0 energy consumption 3× Agent1 |
| Heatmap Texture | Uniform red | Gradient sparse | Clearly asymmetric |
Division of Labor Interpretation:
- Agent0 (Dominant): All neurons highly active, bearing main pushing task and continuous decision-making
- Agent1 (Auxiliary): Only some neurons active, bearing auxiliary tasks (balance monitoring, emergency response, symbol signal reception)
Communication Pattern: “Broadcast-Listen” Consciousness Interaction
┌─────────────────┐
│ Agent0 │
│ "Tonic Consciousness"│
│ Continuous Broadcast│
└────────┬────────┘
│
Symbol Layer Communication
│
┌────────▼────────┐
│ Agent1 │
│ "Gradient Consciousness"│
│ Selective Listening│
│ On-Demand Response│
└─────────────────┘
5. Phase Three: Late Directional Reinforcement (gen=1000)
5.1 Agent0 Activity Pattern
| Metric | Value | Change (vs gen=700) |
|---|---|---|
| Mean Spike Count | 3.95 | Stable |
| Activation Ratio | 100% | Stable |
| Total Spike Moving Average | 1.80-1.98 | Stable |
| Heatmap Feature | Fully red | Stable |
Key Observations:
- Agent0 maintains extremely high activity, role completely solidified
- No significant change, indicating stable dominant role
5.2 Agent1 Activity Pattern
| Metric | Value | Change (vs gen=700) |
|---|---|---|
| Mean Spike Count | 0-2.5 (steeper gradient) | Significantly decreased |
| Activation Ratio | 0-60% (steeper gradient) | Significantly decreased |
| Total Spike Moving Average | ~1.28 | Maintains low level |
| Heatmap Feature | Extremely sparse | Only few neurons active |
Key Observations:
- Agent1 activity further decreases
- Only first 3-5 neurons maintain low-level activity (~2.0-2.5)
- Most neurons completely silent (activation ratio 0%)
- Total spike count ≈ 12-15 (55-68% of gen=700, 18-22% of Agent0)
5.3 Consciousness Pattern: “Sparse Consciousness”
Neuron 0: ██████░░░░░░░░░░░░░░ (2.0) ← Only few active
Neuron 1: █████░░░░░░░░░░░░░░░ (1.8)
Neuron 2: ████░░░░░░░░░░░░░░░░ (1.5)
Neuron 3: ██░░░░░░░░░░░░░░░░░░ (1.0)
Neuron 4: ░░░░░░░░░░░░░░░░░░░░ (0.5)
...
Neuron 15: ░░░░░░░░░░░░░░░░░░░░ (0.0)
Consciousness Characteristics:
- Minimized Consciousness Breadth: Only 3-4 neurons remain active
- “Minimal Consciousness” State: Similar to biological “minimum energy monitoring mode” — retaining only essential sensory and response functions
- Energy Priority: From gen=700 to 1000, Agent1 proactively “shuts down” most neurons, entering extreme minimalist mode
5.4 Division of Labor Analysis (gen=1000)
| Dimension | Agent0 | Agent1 | Interpretation |
|---|---|---|---|
| Activity Intensity | Extremely high (~67) | Extremely low (~12-15) | Agent1 highly energy-efficient |
| Neuronal Differentiation | None | Clear gradient | Agent1 retains only core functions |
| Role Stability | Stable | Continually streamlined | Deepening division of labor |
6. Phase Four: Mature Phase (gen=1500)
6.1 Agent0 Activity Pattern
| Metric | Value | Change (vs gen=1000) |
|---|---|---|
| Mean Spike Count | 3.95 | Stable |
| Activation Ratio | 100% | Stable |
| Total Spike Moving Average | 1.80-1.98 | Stable |
| Heatmap Feature | Fully red | Stable |
Key Observations:
- Agent0 maintains extremely high activity, no degradation
- Indicates dominant role completely solidified
6.2 Agent1 Activity Pattern
| Metric | Value | Change (vs gen=1000) |
|---|---|---|
| Mean Spike Count | 0.3-2.5 (gradient) | Stable |
| Activation Ratio | 0-60% (gradient) | Stable |
| Total Spike Moving Average | ~1.28 | Stable |
| Heatmap Feature | Sparse | Stable |
Key Observations:
- Agent1 activity stabilizes after gen=1000
- Forms stable sparse activation pattern
- Total spike count ≈ 22 (slight rebound from gen=1000)
6.3 Consciousness Pattern: “Contemplative Consciousness”
Neuron 0: ████████░░░░░░░░░░░░ (2.3)
Neuron 1: ███████░░░░░░░░░░░░░ (2.1)
Neuron 2: ██████░░░░░░░░░░░░░░ (1.9)
Neuron 3: █████░░░░░░░░░░░░░░░ (1.7)
Neuron 4: ████░░░░░░░░░░░░░░░░ (1.4)
Neuron 5: ███░░░░░░░░░░░░░░░░░ (1.1)
...
Neuron 15: ░░░░░░░░░░░░░░░░░░░░ (0.3)
Consciousness Characteristics:
- Consciousness Breadth Rebound: 5-8 neurons recover activity (increase from gen=1000)
- Stabilized Consciousness Depth: Gradient structure solidified, forming stable hierarchical information processing
- “Contemplative” State: Similar to meditative “awareness without intervention” — monitoring environment without active intervention, responding only when necessary
6.4 Division of Labor Analysis (gen=1500)
| Dimension | Agent0 | Agent1 | Interpretation |
|---|---|---|---|
| Activity Intensity | Extremely high (~67) | Low (~22) | Stable asymmetry |
| Neuronal Differentiation | None | Stable gradient | Function solidified |
| Role Stability | Completely solidified | Completely solidified | Mature division of labor |
Final Division of Labor:
- Agent0 (Primary Pusher): All neurons continuously highly active, bearing main pushing task
- Agent1 (Auxiliary/Energy-Saver): Only some neurons sparsely active, bearing auxiliary tasks
7. Consciousness Evolution Trajectory
7.1 Agent0: “Tonic Consciousness” (Constant)
Time →→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→→
Consciousness Breadth: ████████████████████ (max, constant)
Consciousness Depth: ░░░░░░░░░░░░░░░░░░░░ (min, constant)
Stability: ████████████████████ (max, constant)
7.2 Agent1: Consciousness Evolution “U-Curve”
Consciousness Breadth
↑
High│ Agent0 (constant)
│
Mid │ Agent1 (gen=700) ●
│ \
│ \
Low │ ● (gen=1000)
│ \
│ \
Mid │ ● (gen=1500)
│
└────────────────────────────────→ Time
700 1000 1500
7.3 Dual-Agent Consciousness Evolution Timeline
| Generation | Agent0 Consciousness | Agent1 Consciousness | Overall Consciousness Mode |
|---|---|---|---|
| 50 | Symmetric high activity | Symmetric high activity | Dual high-energy mode |
| 700 | “Tonic Consciousness” | “Gradient Consciousness” | Asymmetric formation |
| 1000 | “Tonic Consciousness” | “Sparse Consciousness” | Deepening asymmetry |
| 1500 | “Tonic Consciousness” | “Contemplative Consciousness” | Efficient complementarity |
8. Symbol Layer Communication Evolution
8.1 Symbol Layer Activity Comparison
The symbol layer (last 4 neurons) serves as the communication channel between agents:
| Generation | Agent0 Symbol Layer | Agent1 Symbol Layer | Communication Mode |
|---|---|---|---|
| 700 | Continuous high pulse (3.95) | Gradient pulse (0.5-2.0) | Agent0 continuous broadcast, Agent1 selective reception |
| 1000 | Continuous high pulse (3.95) | Extremely low pulse (0-0.5) | Agent0 continuous broadcast, Agent1 nearly no response |
| 1500 | Continuous high pulse (3.95) | Moderate pulse (0.5-1.5) | Stable unidirectional communication established |
8.2 Communication Evolution
| Phase | Agent0 → Symbol | Agent1 ← Symbol | Communication Pattern |
|---|---|---|---|
| Early (700) | Continuous broadcast | Selective reception | Unidirectional, selective |
| Mid (1000) | Continuous broadcast | Minimal reception | Near-zero response |
| Late (1500) | Continuous broadcast | Moderate reception | Stable unidirectional |
9. Consciousness-Level Evolution Patterns
9.1 Relationship Between Division of Labor and Consciousness State
Division of Labor
↑
High│ ● (1500)
│ / \
│ / \
│ / \
│ / \
│ / \
Mid │ ● (1000) \
│ / \
│ / \
│ / \
│ / \
Low │ ● (700) \
│ \
└────────────────────────────────→ Consciousness Asymmetry
Low Mid High
Pattern: The clearer the division of labor, the more asymmetric the consciousness states.
9.2 Consciousness as Adaptive Evolution
| Evolutionary Phase | Agent0 Consciousness | Agent1 Consciousness | Cognitive Strategy |
|---|---|---|---|
| Exploration (50) | High exploration | High exploration | Symmetric exploration |
| Differentiation (700) | Maintains high | Attempts silence | Optimal division search |
| Deepening (1000) | “Tonic” | “Minimal” | Extreme energy saving |
| Maturation (1500) | “Tonic” | “Contemplative” | Efficient complementarity |
9.3 Biological Analogy: Sympathetic vs Parasympathetic
| System | Agent0 | Agent1 |
|---|---|---|
| Analogy | Sympathetic Nervous System | Parasympathetic Nervous System |
| Function | “Fight or flight” (active pushing) | “Rest and digest” (balance monitoring) |
| Activity Level | Continuously high | Silent, on-demand activation |
| Response | Active intervention | Passive response |
10. Conclusions and Implications
10.1 Core Findings
-
Spontaneous Division of Labor Emergence: Without predefined division, two agents spontaneously form “primary pusher-auxiliary” roles
-
Two Consciousness Modes:
- Agent0: “Tonic Consciousness” — continuous high output, automated execution
- Agent1: “Contemplative Consciousness” — gradient sparse activity, monitoring with on-demand response
-
Asymmetric Energy Distribution: Agent0 energy consumption 3-5× Agent1, forming complementary collaboration
-
Consciousness Plasticity: Agent1 undergoes “medium → minimal → stable” consciousness breadth evolution, demonstrating neuronal adaptive capacity
-
Unidirectional Communication Evolution: From bidirectional saturated communication to one-way broadcast-selective listening mode
-
Consciousness Asymmetry as Foundation for Efficient Collaboration: One actively executes, one silently monitors, forming complementarity
10.2 Training Evaluation
- Successful: Directional reinforcement successfully guided efficient complementary division of labor
- Emergent Intelligence: Agents autonomously discovered the superior strategy of “one pushes, one assists”
- Efficiency Improvement: Overall energy consumption reduced by approximately 30-40% (compared to dual high-activity mode)
10.3 Theoretical Significance
-
Observability of Artificial Consciousness: This study demonstrates how to quantify “artificial consciousness states” through neuronal activity patterns
-
Co-evolution of Division of Labor and Consciousness: Role division and consciousness states mutually shape each other, forming stable collaboration patterns
-
Energy-Conscious Cognition: Agents can learn “when to think, when to be silent,” achieving cognitive energy savings
10.4 Future Research Directions
-
Neuron-Action Correlation Analysis: Identify which neurons are specifically responsible for right-push decisions
-
Consciousness State Switching: Can Agent1 actively switch to “high consciousness” mode when necessary?
-
Division of Labor Stability Testing: Verify if different initial conditions produce the same division of labor
-
Three-Agent Consciousness Ecology: Can multiple agents form more complex consciousness interaction networks?
-
Consciousness Interpretability: Can gradient structures be decoded to reveal specific functions (balance, thrust, angle, etc.)?
-
Symbol Communication Decoding: Analyze when Agent1’s symbol layer activates to understand collaboration protocols
-
Division of Labor Reward: Actively encouraging asymmetric strategies may accelerate convergence
11. Appendix: Data Confirmation
11.1 Valid Data Files
| File | Exists | Quality |
|---|---|---|
| consciousness_agent0_gen700_ep0.txt.png | ✓ | Good |
| consciousness_agent0_gen1000_ep0.txt.png | ✓ | Good |
| consciousness_agent0_gen1500_ep0.txt.png | ✓ | Good |
| consciousness_agent1_gen700_ep0.txt.png | ✓ | Good |
| consciousness_agent1_gen1000_ep0.txt.png | ✓ | Good |
| consciousness_agent1_gen1500_ep0.txt.png | ✓ | Good |
11.2 Key Values Confirmation
Agent0 (gen=1500):
- Mean Spike Count: 3.95
- Activation Ratio: 1.00
- Total Spike Moving Average: 1.80-1.98
Agent1 (gen=1500):
- Mean Spike Count Range: 0.3-2.5
- Activation Ratio Range: 0-60%
- Total Spike Moving Average: ~1.28
意识图
Report Completed: March 26, 2026






演示程序
c++代码
best_model_final_v9_1.txt
2.64097
-4.2671
-1.51756
3.08342
1.47481
-1.31492
0.561204
-3.33686
0.970364
4.65745
-3.2742
4.9997
3.8298
0.912659
3.10058
4.99446
4.35861
-5
-4.37312
5
3.15371
-0.422951
-1.21235
2.77163
-4.98611
4.60843
-4.02966
5
3.42434
-0.474388
0.0723107
-1.2951
4.69918
-4.86903
-0.896957
1.40768
-4.99235
3.7869
1.90601
4.53797
3.22383
3.11245
2.2672
-3.29318
2.16567
4.1003
-5
-0.826191
-0.629348
-0.818605
2.43659
-3.98381
4.97382
2.70572
-2.57362
-3.55753
4.6005
-4.95019
-3.69982
0.187319
-2.94814
0.843598
4.45439
-3.93191
4.75332
-1.51723
-5
4.74127
3.50981
-4.55486
-4.4751
-2.0166
5
0.721719
2.60846
1.49439
3.62977
-2.09869
5
2.04447
0.763911
-0.498845
4.17415
-0.674533
-4.9896
3.88896
3.67144
-3.52902
-3.54462
5
3.97616
4.98821
-1.72015
3.6136
1.20243
1.94287
0.842032
-2.08656
-3.98646
-3.07164
-4.94686
-2.83859
-2.98043
-2.8995
1.55893
4.50214
5
0.593955
2.32659
0.125987
3.88362
-1.33918
-3.67296
-2.94529
-2.90254
-4.19969
-0.833069
-1.88
0.508792
-3.51082
-2.58172
-0.723812
0.340893
-0.0809199
-2.1482
-5
2.5685
-5
-2.41191
-3.03113
2.37595
-3.0186
3.58537
3.60022
-1.67895
4.32071
5
-2.30862
-4.22208
4.08969
3.98397
4.14802
3.43294
-2.01614
-3.41794
4.99271
3.65886
-1.67524
-3.42894
-4.26027
0.56273
-1.60134
-2.69251
2.23105
3.62026
-2.60787
0.421456
-4.33835
-1.29767
4.24975
-5
3.59844
-1.3036
-0.134036
-1.54101
-1.56821
3.35252
1.95306
0.612443
1.284
3.68503
-2.92217
-5
4.98836
-1.47461
-4.87529
-1.5237
-1.24759
3.23401
2.47798
-1.34649
-0.567147
5
1.66411
-2.61348
3.09924
-3.02555
-3.33979
5
2.33553
4.19723
-4.97955
0.873651
2.50536
-3.88709
-1.21528
-2.75376
-0.938089
5
-4.54234
-1.79352
1.4345
3.09558
2.89292
-2.34668
3.05002
2.13581
-1.42648
-0.921491
4.93969
4.44464
-2.6954
0.137355
-1.70593
0.858039
0.357402
4.3803
-3.12255
-5
-3.24484
-0.29707
-5
-2.74009
-2.46768
4.66506
2.11378
2.65259
-1.85901
4.99929
2.17916
0.231185
1.563
1.25256
1.69574
0.594262
-4.34175
-2.02003
-4.89902
-4.80821
0.0035492
-4.54155
-1.11736
2.72803
3.69715
3.04575
-2.14417
0.482102
5
4.78198
4.03876
1.06147
-5
-4.95415
-1.98134
0.0475862
-2.41932
-0.0593832
-0.649927
-1.44983
-4.99282
5
-1.57941
4.86642
-3.50047
1.43832
-5
-0.544199
-3.22094
0.0574737
-3.92391
-0.892034
5
4.42994
3.21593
5
4.10736
4.7863
-4.98304
-2.51028
1.16219
3.92291
-3.35726
-0.570141
-0.951276
3.64192
3.7888
-1.72404
3.29126
1.59177
4.31848
-3.88176
1.7963
4.79086
-4.6444
0.341975
-0.947259
-1.92332
-2.63007
-4.87019
-4.94563
0.411629
-4.40184
3.07399
-2.16216
-4.8
3.1772
3.14586
-4.97523
-1.6473
-3.23216
2.97216
5
-1.66918
1.12986
-1.04735
2.15213
-5
-4.93627
-4.07335
-4.2454
-1.49523
-5
3.7617
-3.43378
-0.828722
-0.392973
-0.87575
-1.75172
4.7491
2.6294
5
-4.83546
-3.33843
2.77175
3.98852
4.58201
2.90604
-0.682612
1.68529
-2.2433
0.388984
-0.191372
2.15749
-4.21189
-1.08521
4.05768
-0.357764
-3.38782
-0.124778
0.426328
0.778225
3.45836
-4.13837
-3.76842
-2.47217
3.76994
-4.96852
-2.93037
1.01142
-2.27193
-2.46618
-4.98613
3.33999
3.39028
-4.99005
-3.6338
2.4388
1.39645
-0.397231
-5
4.8869
-4.48288
0.83333
0.849891
-1.93127
4.50236
4.08255
3.07119
0.15228
-4.15539
-3.36303
-2.0801
-4.80067
4.92982
-3.64204
-1.94489
2.88125
4.31422
-5
-2.37112
-4.58252
-3.83718
3.46609
-0.668855
-1.7725
-2.4746
0.685944
-3.3519
-4.48594
5
-1.35382
-3.27288
-5
-3.05549
-5
-1.96512
-3.27244
2.06984
0.125398
4.73904
3.17064
4.33891
5
-2.1118
-0.106686
-4.98854
3.9289
-0.478979
0.704632
1.13502
-1.84937
-3.6945
-4.31248
-1.89121
5
2.20737
-4.28751
-2.85736
3.65943
-5
/**
* 演示程序:加载训练好的 V9 模型,在控制台动态显示多智能体推箱过程。
* 编译:g++ -O3 -std=c++17 demo_final_v9.cpp -o demo_final_v9
* 运行:./demo_final_v9 [模型文件路径] [演示局数]
* 默认模型文件:best_model_final_v9.txt,演示 1 局。
*/
#include <iostream>
#include <vector>
#include <cmath>
#include <random>
#include <chrono>
#include <fstream>
#include <string>
#include <thread>
#include <unistd.h> // for usleep
using namespace std;
// ==================== 环境常量 ====================
constexpr double GRAVITY = 9.8;
constexpr double MASSCART = 1.0;
constexpr double MASSPOLE = 0.1;
constexpr double LENGTH = 0.5;
constexpr double FORCE_MAG = 10.0;
constexpr double TAU = 0.02;
constexpr double FOURTHIRDS = 4.0/3.0;
constexpr double POSITION_LIMIT = 2.4;
constexpr double ANGLE_LIMIT = 12.0 * M_PI / 180.0;
constexpr int MAX_STEPS = 1000;
constexpr int NUM_AGENTS = 2;
// ==================== SNN 参数 ====================
constexpr int D = 4;
constexpr double V_REST = -70.0;
constexpr double V_RESET = -75.0;
constexpr double V_THRESH_BASE = -55.0;
constexpr double THRESH_INTERVAL = 1.0;
constexpr double TAU_M = 10.0;
constexpr double R_M = 10.0;
constexpr double DT = 1.0;
constexpr int N_INPUT = 4;
constexpr int N_HIDDEN1 = 8;
constexpr int N_HIDDEN2 = 5;
constexpr int N_SYMBOL = 4;
constexpr int N_OUTPUT = 2;
constexpr int N_EXT_CHANNELS = (NUM_AGENTS - 1) * N_SYMBOL;
// 基因维度(与训练一致)
constexpr int DIM_SINGLE = N_HIDDEN1 * N_INPUT
+ N_HIDDEN1 * N_HIDDEN1
+ N_HIDDEN2 * N_HIDDEN1
+ N_SYMBOL * N_HIDDEN2
+ N_OUTPUT * N_SYMBOL
+ N_HIDDEN1 * N_EXT_CHANNELS
+ N_HIDDEN1
+ N_HIDDEN2
+ N_SYMBOL
+ N_OUTPUT;
constexpr int DIM_TOTAL = NUM_AGENTS * DIM_SINGLE;
constexpr double W_MAX = 5.0;
// 随机数生成器
mt19937 rng(chrono::steady_clock::now().time_since_epoch().count());
uniform_real_distribution<double> uniform_01(0.0, 1.0);
// ==================== MSF 神经元 ====================
struct MSFNeuron {
vector<double> thresholds;
MSFNeuron() {
thresholds.resize(D);
for (int d = 0; d < D; ++d) thresholds[d] = V_THRESH_BASE + d * THRESH_INTERVAL;
}
int step(double I_ext) {
double v = V_REST + R_M * I_ext;
if (v > V_THRESH_BASE + (D-1)*THRESH_INTERVAL + 10)
v = V_THRESH_BASE + (D-1)*THRESH_INTERVAL + 10;
int spike_count = 0;
for (int d = 0; d < D; ++d) {
if (v >= thresholds[d]) spike_count++;
else break;
}
return spike_count;
}
};
// ==================== 带语言符号的循环 SNN(仅前向) ====================
class SocialSNN {
private:
vector<double> w_in_h1, w_fb_h1, w_h1_h2, w_h2_sym, w_sym_out, w_ext_h1;
vector<double> bias_h1, bias_h2, bias_sym, bias_out;
vector<MSFNeuron> hidden1, hidden2, symbol, output;
vector<int> prev_spike_h1, prev_spike_h2, prev_spike_sym;
vector<double> ext_current;
public:
SocialSNN(const vector<double>& genes, size_t offset) {
hidden1.resize(N_HIDDEN1); hidden2.resize(N_HIDDEN2);
symbol.resize(N_SYMBOL); output.resize(N_OUTPUT);
ext_current.assign(N_HIDDEN1, 0.0);
size_t pos = offset;
w_in_h1.resize(N_HIDDEN1 * N_INPUT);
w_fb_h1.resize(N_HIDDEN1 * N_HIDDEN1);
w_h1_h2.resize(N_HIDDEN2 * N_HIDDEN1);
w_h2_sym.resize(N_SYMBOL * N_HIDDEN2);
w_sym_out.resize(N_OUTPUT * N_SYMBOL);
w_ext_h1.resize(N_HIDDEN1 * N_EXT_CHANNELS);
bias_h1.resize(N_HIDDEN1); bias_h2.resize(N_HIDDEN2);
bias_sym.resize(N_SYMBOL); bias_out.resize(N_OUTPUT);
for (size_t i = 0; i < w_in_h1.size(); ++i) w_in_h1[i] = genes[pos++];
for (size_t i = 0; i < w_fb_h1.size(); ++i) w_fb_h1[i] = genes[pos++];
for (size_t i = 0; i < w_h1_h2.size(); ++i) w_h1_h2[i] = genes[pos++];
for (size_t i = 0; i < w_h2_sym.size(); ++i) w_h2_sym[i] = genes[pos++];
for (size_t i = 0; i < w_sym_out.size(); ++i) w_sym_out[i] = genes[pos++];
for (size_t i = 0; i < w_ext_h1.size(); ++i) w_ext_h1[i] = genes[pos++];
for (int i = 0; i < N_HIDDEN1; ++i) bias_h1[i] = genes[pos++];
for (int i = 0; i < N_HIDDEN2; ++i) bias_h2[i] = genes[pos++];
for (int i = 0; i < N_SYMBOL; ++i) bias_sym[i] = genes[pos++];
for (int i = 0; i < N_OUTPUT; ++i) bias_out[i] = genes[pos++];
reset_state();
}
void reset_state() {
prev_spike_h1.assign(N_HIDDEN1, 0);
prev_spike_h2.assign(N_HIDDEN2, 0);
prev_spike_sym.assign(N_SYMBOL, 0);
fill(ext_current.begin(), ext_current.end(), 0.0);
}
void set_external_current(const vector<double>& ext) { ext_current = ext; }
int forward(const vector<double>& input) {
vector<double> I_h1(N_HIDDEN1);
for (int i = 0; i < N_HIDDEN1; ++i) {
I_h1[i] = bias_h1[i];
for (int j = 0; j < N_INPUT; ++j)
I_h1[i] += w_in_h1[i * N_INPUT + j] * input[j];
for (int j = 0; j < N_HIDDEN1; ++j)
I_h1[i] += w_fb_h1[i * N_HIDDEN1 + j] * prev_spike_h1[j];
I_h1[i] += ext_current[i];
}
vector<int> spike_h1(N_HIDDEN1);
for (int i = 0; i < N_HIDDEN1; ++i) spike_h1[i] = hidden1[i].step(I_h1[i]);
vector<double> I_h2(N_HIDDEN2);
for (int i = 0; i < N_HIDDEN2; ++i) {
I_h2[i] = bias_h2[i];
for (int j = 0; j < N_HIDDEN1; ++j)
I_h2[i] += w_h1_h2[i * N_HIDDEN1 + j] * spike_h1[j];
}
vector<int> spike_h2(N_HIDDEN2);
for (int i = 0; i < N_HIDDEN2; ++i) spike_h2[i] = hidden2[i].step(I_h2[i]);
vector<double> I_sym(N_SYMBOL);
for (int i = 0; i < N_SYMBOL; ++i) {
I_sym[i] = bias_sym[i];
for (int j = 0; j < N_HIDDEN2; ++j)
I_sym[i] += w_h2_sym[i * N_HIDDEN2 + j] * spike_h2[j];
}
vector<int> spike_sym(N_SYMBOL);
for (int i = 0; i < N_SYMBOL; ++i) spike_sym[i] = symbol[i].step(I_sym[i]);
vector<double> I_out(N_OUTPUT);
for (int i = 0; i < N_OUTPUT; ++i) {
I_out[i] = bias_out[i];
for (int j = 0; j < N_SYMBOL; ++j)
I_out[i] += w_sym_out[i * N_SYMBOL + j] * spike_sym[j];
}
vector<int> out_spikes(N_OUTPUT);
for (int i = 0; i < N_OUTPUT; ++i) out_spikes[i] = output[i].step(I_out[i]);
int action = 0;
int max_spike = out_spikes[0];
for (int i = 1; i < N_OUTPUT; ++i) {
if (out_spikes[i] > max_spike) {
max_spike = out_spikes[i];
action = i;
}
}
prev_spike_h1 = spike_h1;
prev_spike_h2 = spike_h2;
prev_spike_sym = spike_sym;
fill(ext_current.begin(), ext_current.end(), 0.0);
return action;
}
vector<int> get_symbol_spikes() const { return prev_spike_sym; }
};
// ==================== 倒立摆小车 ====================
class CartPole {
private:
double x, x_dot, theta, theta_dot;
int steps;
public:
CartPole() { reset(); }
void reset() {
uniform_real_distribution<double> dist_pos(-0.05, 0.05);
uniform_real_distribution<double> dist_angle(-0.05, 0.05);
x = dist_pos(rng);
x_dot = dist_pos(rng) * 0.1;
theta = dist_angle(rng);
theta_dot = dist_angle(rng) * 0.1;
steps = 0;
}
void update(double force) {
double total_mass = MASSCART + MASSPOLE;
double mass_pole_len = MASSPOLE * LENGTH;
double sin_theta = sin(theta);
double cos_theta = cos(theta);
double temp = (force + mass_pole_len * theta_dot * theta_dot * sin_theta) / total_mass;
double theta_acc = (GRAVITY * sin_theta - cos_theta * temp) /
(LENGTH * (FOURTHIRDS - (MASSPOLE * cos_theta * cos_theta) / total_mass));
double x_acc = temp - (mass_pole_len * theta_acc * cos_theta) / total_mass;
x += TAU * x_dot;
x_dot += TAU * x_acc;
theta += TAU * theta_dot;
theta_dot += TAU * theta_acc;
steps++;
}
vector<double> get_state() const {
vector<double> state(N_INPUT);
state[0] = x / POSITION_LIMIT;
double xd_clip = max(-3.0, min(3.0, x_dot));
state[1] = xd_clip / 3.0;
state[2] = theta / ANGLE_LIMIT;
double td_clip = max(-2.0, min(2.0, theta_dot));
state[3] = td_clip / 2.0;
return state;
}
bool is_done() const {
return (x < -POSITION_LIMIT || x > POSITION_LIMIT ||
theta < -ANGLE_LIMIT || theta > ANGLE_LIMIT ||
steps >= MAX_STEPS);
}
int get_steps() const { return steps; }
double get_x() const { return x; }
double get_angle() const { return theta; }
};
// ==================== 多智能体协作环境 ====================
class MultiCargoEnv {
private:
vector<CartPole> carts;
double cargo_x, cargo_v;
int steps;
double cargo_mass;
double goal_x;
public:
MultiCargoEnv(int n_agents) : carts(n_agents), cargo_x(0.0), cargo_v(0.0), steps(0),
cargo_mass(3.0), goal_x(5.0) {}
void reset() {
for (auto& c : carts) c.reset();
cargo_x = 0.0;
cargo_v = 0.0;
steps = 0;
}
void step(const vector<int>& actions) {
double total_force = 0.0;
for (size_t i = 0; i < carts.size(); ++i) {
double force = (actions[i] == 0) ? -FORCE_MAG : FORCE_MAG;
carts[i].update(force);
total_force += force;
}
double cargo_acc = total_force / cargo_mass;
cargo_v += cargo_acc * TAU;
cargo_x += cargo_v * TAU;
steps++;
}
bool is_done() const {
for (const auto& c : carts)
if (c.is_done()) return true;
return (cargo_x < -5.0 || cargo_x > 5.0 || steps >= MAX_STEPS);
}
vector<vector<double>> get_observations() const {
vector<vector<double>> obs;
for (const auto& c : carts) obs.push_back(c.get_state());
return obs;
}
double get_cargo_x() const { return cargo_x; }
int get_steps() const { return steps; }
const vector<CartPole>& get_carts() const { return carts; }
};
// ==================== 加载模型 ====================
bool load_parameters(vector<double>& params, const string& filename) {
ifstream fin(filename);
if (!fin.is_open()) return false;
params.clear();
double val;
while (fin >> val) params.push_back(val);
return params.size() == DIM_TOTAL;
}
// ==================== 演示主函数 ====================
void display_status(const MultiCargoEnv& env, int step) {
// 清屏并移动光标到左上角
cout << "\033[2J\033[H";
const double x_range = 10.0; // 显示范围 -5..5
const int width = 70; // 字符宽度
const double scale = width / x_range;
// 显示货物位置(用 'C' 表示)
double cargo_x = env.get_cargo_x();
int cargo_pos = static_cast<int>((cargo_x + x_range/2) * scale);
cargo_pos = max(0, min(width-1, cargo_pos));
// 显示每个小车位置(用 '1' 和 '2' 表示)
const auto& carts = env.get_carts();
vector<int> cart_positions(NUM_AGENTS);
for (int i = 0; i < NUM_AGENTS; ++i) {
double x = carts[i].get_x();
int pos = static_cast<int>((x + x_range/2) * scale);
cart_positions[i] = max(0, min(width-1, pos));
}
// 绘制轨道
string line(width, '-');
line[cargo_pos] = 'C';
for (int i = 0; i < NUM_AGENTS; ++i) {
if (line[cart_positions[i]] == 'C')
line[cart_positions[i]] = 'M'; // 货物和小车重叠
else if (line[cart_positions[i]] != '-')
line[cart_positions[i]] = 'X'; // 多个小车重叠
else
line[cart_positions[i]] = (i == 0 ? '1' : '2');
}
cout << "Step: " << step << "/" << MAX_STEPS << "\n";
cout << "Cargo X: " << cargo_x << "\n";
for (int i = 0; i < NUM_AGENTS; ++i) {
cout << "Cart " << i+1 << " X: " << carts[i].get_x()
<< ", Angle: " << carts[i].get_angle() * 180/M_PI << " deg\n";
}
cout << "\n";
cout << "|" << line << "|\n";
cout << " -5.0 0.0 5.0\n";
cout.flush();
}
int main(int argc, char* argv[]) {
string model_file = "best_model_final_v9.txt";
int num_episodes = 1;
if (argc > 1) model_file = argv[1];
if (argc > 2) num_episodes = stoi(argv[2]);
// 加载模型
vector<double> genes;
if (!load_parameters(genes, model_file)) {
cerr << "错误:无法加载模型文件 " << model_file << ",请确保文件存在且参数数量正确。\n";
return 1;
}
cout << "模型加载成功,参数个数:" << genes.size() << endl;
// 演示多局
for (int ep = 0; ep < num_episodes; ++ep) {
MultiCargoEnv env(NUM_AGENTS);
vector<SocialSNN> agents;
for (int i = 0; i < NUM_AGENTS; ++i) {
size_t offset = i * DIM_SINGLE;
agents.emplace_back(genes, offset);
}
for (auto& a : agents) a.reset_state();
vector<vector<int>> prev_symbols(NUM_AGENTS, vector<int>(N_SYMBOL, 0));
int step = 0;
while (!env.is_done()) {
auto obs = env.get_observations();
// 计算外部输入(符号)
vector<vector<double>> ext_currents(NUM_AGENTS, vector<double>(N_HIDDEN1, 0.0));
for (int i = 0; i < NUM_AGENTS; ++i) {
int chan_offset = 0;
for (int j = 0; j < NUM_AGENTS; ++j) {
if (i == j) continue;
for (int k = 0; k < N_SYMBOL; ++k) {
int ext_idx = chan_offset + k;
if (ext_idx < N_EXT_CHANNELS) ext_currents[i][ext_idx] += prev_symbols[j][k];
}
chan_offset += N_SYMBOL;
}
}
vector<int> actions(NUM_AGENTS);
for (int i = 0; i < NUM_AGENTS; ++i) {
agents[i].set_external_current(ext_currents[i]);
actions[i] = agents[i].forward(obs[i]);
}
// 更新符号
for (int i = 0; i < NUM_AGENTS; ++i) prev_symbols[i] = agents[i].get_symbol_spikes();
env.step(actions);
step++;
// 显示状态(每秒约 50 帧,可根据需要调整)
display_status(env, step);
usleep(20000); // 20ms 刷新一次,约 50fps
}
cout << "\n第 " << ep+1 << " 局结束,总步数:" << env.get_steps() << endl;
if (ep + 1 < num_episodes) {
cout << "按回车键继续下一局...";
cin.get();
}
}
return 0;
}

Display Interface Explanation
Example Screen
Step: 150/1000
Cargo X: 2.35
Cart 1 X: 2.30, Angle: 1.5 deg
Cart 2 X: -0.80, Angle: -0.5 deg
|---------------------1----M-------------------------|
-5.0 0.0 5.0
The Track Line
The bottom line represents a horizontal track ranging from -5.0 to 5.0:
Position: -5.0 .......... 0.0 .......... 5.0
Track: |-----------------------------------------------|
Each character represents a position on the track:
| Symbol | Meaning |
|---|---|
- | Empty track, nothing here |
1 | Agent 0 (Cart 1) is here |
2 | Agent 1 (Cart 2) is here |
C | Cargo is here, no cart touching it |
M | Merge - Cargo and a cart overlap (being pushed) |
X | Two carts overlap (collision) |
Examples
Case 1: Cart 1 alone
|---------------------1------------------------------|
→ Cart 1 is at position ~1.2, cargo is elsewhere
Case 2: Cart 1 pushing cargo
|---------------------M------------------------------|
→ M means Cart 1 and cargo are at the same position — pushing is happening
Case 3: Both carts pushing cargo
|---------------------M------------------------------|
→ If both carts overlap with cargo, it still shows M
Case 4: Two carts colliding
|---------------------X------------------------------|
→ X means Cart 1 and Cart 2 are at the same position, but cargo is elsewhere
Case 5: Cargo being pushed, other cart elsewhere
|---------------------M-------2----------------------|
→ M position: Cart 1 is pushing cargo
→ 2 position: Cart 2 is far away, not participating in pushing
Quick Guide
What to look for:
| Question | Look for |
|---|---|
| Who is pushing the cargo? | M shows where the cargo is being pushed |
| Where is the cargo? | M (if being pushed) or C (if free) |
| Where is Cart 1? | 1 (or M/X if overlapping) |
| Where is Cart 2? | 2 (or M/X if overlapping) |
| Are carts colliding? | X |
Summary
M= Pushing action happening here1/2= Individual cart positionsC= Cargo alone (no cart touching)X= Carts colliding-= Empty track
This visualization lets you see at a glance:
- Which cart is pushing the cargo (
M) - Whether both carts are cooperating at the same spot (
M) - Whether they are separated (
1and2in different places) - Whether they accidentally collided (
X)
更多推荐


所有评论(0)