Alignment & Preference Tuning · Emerging · Advanced
Process-Supervised Reward Models
Also known as: PRMs, Step-Level Feedback
Reward models trained to evaluate and provide scalar feedback on individual reasoning steps rather than just the final output.
What Process-Supervised Reward Models is
Process-Supervised Reward Models (PRMs) significantly improve accuracy on complex math and logical reasoning benchmarks.
How it works
Human or automated annotators grade each step in a chain of thought, guiding the model to favor sound logical progression.
Why it matters
Prevents models from reaching correct answers through flawed or hallucinated reasoning steps.
Common uses
- →Mathematical reasoning models
- →Code synthesis verification
- →Logical problem solving
More in this collection
Browse all AI Concepts