Latent-GRPO: Reinforcement Learning in Continuous Thought Space
DEV Community

Latent-GRPO: Reinforcement Learning in Continuous Thought Space

Latent-GRPO: Reinforcement Learning in Continuous Thought Space

When you sit down to solve a complex puzzle or plan three moves ahead in chess, do you narrate every synaptic firing to yourself in full, grammatically correct English sentences? Of course not. Human cognition operates across continuous, multi-dimensional mental representations-spatial, relational, intuitive. We only convert our thoughts into human language when we need to speak to someone else. Yet the dominant paradigm in AI reasoning today (from DeepSeek-R1 to OpenAI o1) forces neural networks to do the exact opposite. We train models to emit thousands of discrete text tokens enclosed in `

Read on DEV Community ↗ ← Back to News

Comments

No comments yet. Start the discussion.