Santa Fe
Institute
  • Research
    • Themes
    • Projects
    • SFI Press
    • Researchers
    • Publications
    • Library
    • Sponsored Research
    • Fellowships
    • Miller Scholarships
  • News + Events
    • News
    • Newsletters
    • Podcasts
    • SFI in the Media
    • Media Center
    • Events
    • Community
    • Journalism Fellowship
  • Education
    • Programs
    • Projects
    • Alumni
    • Complexity Explorer
    • Education FAQ
    • Postdoctoral Research
    • Education Supporters
  • People
    • Researchers
    • Fractal Faculty
    • Staff
    • Miller Scholars
    • Trustees
    • Governance
    • Resident Artists
    • Research Supporters
  • Applied Complexity
    • Office
    • Applied Projects
    • ACtioN
    • Applied Fellows
    • Studios
    • Applied Events
    • Login
  • Give
    • Give Now
    • Ways to Give
    • Contact
  • About
    • About SFI
    • Engage
    • Complex Systems
    • FAQ
    • Campuses
    • Jobs
    • Contact
    • Library
    • Employee Portal

Science for a Complex World

Events

Here's what's happening

Give

You make SFI possible

Subscribe

Sign up for research news

Connect

Follow us on social media

© 2026 Santa Fe Institute. All rights reserved. This site is supported by the Miller Omega Program.

Home / News

Melanie Mitchell: What does it mean to align AI with human values? (Quanta)

Cyborg, created with DiffusionBee AI
December 19, 2022

In her latest column for Quanta Magazine, SFI Professor Melanie Mitchell considers the implications of a machine learning technique called “Inverse Reinforcement Learning.” Researchers have used the technique to train machines to play video games by observing humans, and do backflips in response to human feedback. By bypassing goal-oriented techniques, like in the famous thought experiment involving a superintelligence tasked with producing paper clips, IRL proponents hope to bring AI into better alignment with human ethics.

“An essential first step toward teaching machines ethical concepts is to enable machines to grasp humanlike concepts in the first place,” writes Mitchell. But "ethical notions such as kindness and good behavior are much more complex and context-dependent than anything IRL has mastered so far.”

Without a better scientific theory of intelligence, we may be ill-equipped to tackle AI’s most important problem.

Read the column, "What Does It Mean to Align AI With Human Values?” In Quanta (December 13, 2022)

EXCERPT

Computers frequently misconstrue what we want them to do, with unexpected and often amusing results. One machine learning researcher, for example, while investigating an image classification program’s suspiciously good results, discovered that it was basing classifications not on the image itself, but on how long it took to access the image file — the images from different classes were stored in databases with slightly different access times. Another enterprising programmer wanted his Roomba vacuum cleaner to stop bumping into furniture, so he connected the Roomba to a neural network that rewarded speed but punished the Roomba when the front bumper collided with something. The machine accommodated these objectives by always driving backward

But the community of AI alignment researchers sees a darker side to these anecdotes. In fact, they believe that the machines’ inability to discern what we really want them to do is an existential risk. To solve this problem, they believe, we must find ways to align AI systems with human preferences, goals and values...





Share
  • Sign Up For SFI News
News Media Contact

Santa Fe Institute

Office of Communications
news@santafe.edu
505-984-8800



  • Tags
  • Opinion


More SFI News

View All News

SFI Professors Give Judges Advice on AI

John Krakauer named director of Champalimaud's Centre for Restorative Neurotechnology

Book Review: "Tipping out of Trouble: How Societies Transformed and How We Can Do So Again"

In Memoriam: Jim Rutt

Does intelligence ‘emerge’ in large language models?

Your dominant hand is made, not born

A bird song almost too quiet to hear

Model redefining conformity excels against real-world data

Decoding animal minds

SFI External Professor Nicholas de Monchaux named Dean of UC Berkeley College of Environmental Design

Simon Levin named Fellow of the Royal Society

Brian Enquist receives Robert H. MacArthur Award

Han van der Maas named director of Amsterdam’s Institute for Advanced Study

Marina Dubova receives Dissertation Prize

Smart parts for smart wholes

Aaron Clauset receives honors from AAAS and University of New Mexico

Laurent Hébert-Dufresne receives Erdős-Rényi Prize

Why noise may be the key to understanding cell group patterns

Reinventing democracy before it breaks

Do deep learning models recognize 3D shapes in the same way humans do?