Moyan AI Training Institution LogoMoyan AI

Alignment & Preference Tuning · Emerging · Advanced

Process-Supervised Reward Models

Also known as: PRMs, Step-Level Feedback

Reward models trained to evaluate and provide scalar feedback on individual reasoning steps rather than just the final output.

What Process-Supervised Reward Models is

Process-Supervised Reward Models (PRMs) significantly improve accuracy on complex math and logical reasoning benchmarks.

How it works

Human or automated annotators grade each step in a chain of thought, guiding the model to favor sound logical progression.

Why it matters

Prevents models from reaching correct answers through flawed or hallucinated reasoning steps.

Common uses

  • Mathematical reasoning models
  • Code synthesis verification
  • Logical problem solving

More in this collection

Browse all AI Concepts