Safety · Fast-moving · Intermediate
AI Safety
The field concerned with ensuring AI systems behave as intended and do not cause harm.
What AI Safety is
AI safety covers present-day harms such as misinformation, bias and misuse, and longer-horizon questions about oversight of increasingly capable systems. Both agendas share techniques: evaluation, interpretability, guardrails and human oversight.
How it works
Work includes capability and dangerous-capability evaluations before release, alignment training, interpretability research, staged deployment, monitoring and incident response.
Why it matters
Safety practice determines whether AI deployment in medicine, finance, infrastructure and education is defensible, and it increasingly carries legal weight.
Common uses
- →Pre-deployment evaluation
- →Deployment policy and staged rollout
- →Model cards and system documentation
Strengths
- ✓Reduces real, measurable harm
- ✓Increasingly required by regulation
Watch for
- ✓Trades off against release speed
- ✓Evaluation science is immature
Continue exploring
More in this collection
Browse all AI ConceptsSources & References
NIST AI Risk Management Framework