New report explores safe oversight of advanced AI systems

Dept of Industry, Science and Resources

The report looks at how highly capable AI systems may be able to influence an overseers’ beliefs, even while correctly performing tasks.

The AI Safety Institute commissioned CSIRO to explore scalable oversight, an aspect of AI alignment. Findings will contribute to international research under the UK AI Security Institute’s Alignment Project.

Scalable oversight looks at how an overseer of AI can reliably and accurately evaluate and guide AI system behaviour as those systems become increasingly capable.

The Epistemic safety in scalable oversight report explores the need for an AI overseer to consider both:

  • whether the AI system it oversees is producing correct responses
  • what other unintended influences its responses might have.

The report introduces a framework and practical benchmarks and metrics for scientifically investigating and evaluating AI outputs. It provides a deeper understanding of scalable oversight and what considerations are needed as this area of research progresses.

This research supports the Australian Government’s National AI Plan. It improves our understanding of emerging AI risks and strengthens Australia’s capability to develop and deploy AI responsibly.

/Public Release. View in full here.