Model G20 2027 at FLAME University — registrations now open

Technology & Engineering

AI Safety & Alignment

Also known as AI alignment, AI risk, AI safety

how to make powerful AI systems safe and aligned with human values — the risks, and the work to keep them under control

Key people

  • Stuart RussellAI safety pioneer
  • Nick BostromExistential risk philosopher
  • Eliezer YudkowskyAlignment theorist
  • Paul ChristianoTechnical alignment researcher
  • Dario AmodeiApplied AI safety leader

Timeline

  • 1942Isaac Asimov's fictional 'Three Laws of Robotics' provided an early, influential framework for considering the ethical control of intelligent machines.
  • 1948Norbert Wiener's 'Cybernetics' established the scientific study of control and communication in complex systems, implicitly raising questions about the autonomous behaviour of future machines.
  • 2000The Singularity Institute for Artificial Intelligence, later MIRI, was founded, marking the formal establishment of an organisation dedicated to the long-term safety and alignment of advanced AI.
  • 2014Nick Bostrom's 'Superintelligence: Paths, Dangers, Strategies' popularised the concept of existential risk from advanced AI, bringing the topic to a wider academic and public audience.
  • 2015OpenAI was founded with a stated mission to ensure artificial general intelligence benefits all of humanity, explicitly prioritising safety and alignment alongside capability development.
  • 2017The Asilomar AI Principles were released, a consensus document signed by thousands of researchers and experts, outlining guidelines for beneficial AI development and deployment.
  • 2023The inaugural AI Safety Summit was held at Bletchley Park, UK, bringing together governments, leading AI companies, and researchers to discuss frontier AI risks and foster international collaboration.

Read

  • The Alignment Problem: Machine Learning and Human ValuesBrian Christian · 2020Book
  • Human Compatible: Artificial Intelligence and the Problem of ControlStuart Russell · 2019Book
  • Superintelligence: Paths, Dangers, StrategiesNick Bostrom · 2014Book
  • The AI Revolution: The Road to SuperintelligenceTim Urban · 2015Paper
  • Concrete Problems in AI SafetyDario Amodei et al. · 2016Paper

Watch

  • AI Safety: The BasicsRobert Miles AI SafetyVideo
  • Do You Trust This Computer?Chris PaineVideo
  • The Age of AIYouTube OriginalsVideo

Voices to follow

  • Stuart RussellWikipediaProfessor of Computer Science, University of California, Berkeley; author of "Human Compatible"
  • Nick BostromOfficial siteProfessor of Philosophy, University of Oxford; Founding Director, Future of Humanity Institute
  • Dario AmodeiLinkedInCo-founder and CEO, Anthropic
  • Geoffrey Hinton@geoffreyhintonEmeritus Professor, University of Toronto; Turing Award laureate

Debates

  • Should we hit the brakes on making AI super smart until we're sure it's safe?One view: Taking more time lets us really understand the risks and fix potential problems before they get too big. · Another: Slowing down means missing out on AI's amazing potential to solve big problems like diseases or climate change.Open question
  • Can we really teach AI to understand and care about what humans think is right or wrong?One view: We can program AI with rules and give it lots of examples of good and bad behavior, just like teaching a child. · Another: Human values are super complex, often contradictory, and change across cultures – too hard for a machine to fully grasp.Open question
  • Should we let AI make big, life-changing decisions, like who gets a loan or what medical treatment someone needs?One view: AI can process huge amounts of information much faster and more accurately than humans, leading to better, fairer decisions. · Another: If AI makes a mistake in a critical decision, who is responsible? It's hard to hold a machine accountable.Open question

Glossary

  • AI AlignmentAI Alignment means making sure that an AI's goals and actions match what humans actually want and value. It's about teaching the AI to understand and follow our intentions, not just its own programming. For example, if you tell a cleaning robot to "clean the house," alignment means it cleans the house thoroughly without accidentally throwing away important papers or breaking fragile items.
  • AI EthicsAI Ethics involves the moral principles and rules that guide how artificial intelligence systems should be designed, used, and developed. It's about making sure AI acts in a way that is fair, responsible, and respects human values. For example, an ethical AI system for hiring would be designed to avoid bias against certain groups and treat all applicants fairly.
  • AI SafetyAI Safety is about making sure powerful artificial intelligence systems don't cause harm or danger to humans. It's like putting safety features on a car to prevent accidents. For example, ensuring a self-driving car's AI is safe means it won't suddenly swerve into traffic or fail to stop at a red light.
  • Bias (in AI)Bias in AI happens when an artificial intelligence system makes unfair or prejudiced decisions because of the data it was trained on. This data might reflect existing human biases or be incomplete. For example, if an AI trained mostly on images of male doctors struggles to recognize female doctors, it shows a bias from the training data.
  • Control Problem (in AI)The Control Problem in AI refers to the challenge of making sure that very powerful artificial intelligence systems remain under human direction and don't act against our wishes or interests. It's about keeping the AI in charge of its tasks, but not in charge of us. For example, if an AI designed to manage a city's traffic decides the most efficient solution is to ban all personal cars, that would be an AI acting outside human control in a problematic way.
  • Guardrails (AI)Guardrails in AI are specific rules, limits, or safety mechanisms put in place to prevent an artificial intelligence system from doing harmful, unethical, or unwanted things. They act like boundaries to keep the AI within safe operating zones. For example, a guardrail for a chatbot might be a rule that prevents it from generating hate speech or giving medical advice.
  • Hallucination (AI)Hallucination in AI happens when an artificial intelligence system, especially one that generates text or images, creates information that sounds real but is actually false or made up. It's like the AI is confidently guessing or inventing things. For example, if you ask a chatbot for sources on a topic and it invents book titles and authors that don't exist, that's an AI hallucination.
  • Red Teaming (AI)Red Teaming in AI is the process of intentionally trying to find weaknesses, flaws, or potential dangers in an artificial intelligence system by testing it in creative and sometimes adversarial ways. It's like playing the role of a hacker to make the system stronger. For example, a team might try to trick an AI image generator into creating inappropriate content to identify and fix its vulnerabilities before it's released to the public.
  • Transparency (in AI)Transparency in AI means being able to understand how an artificial intelligence system makes its decisions or comes up with its answers. It's like being able to see inside a black box to know why something happened. For example, if an AI recommends a certain movie, transparency would mean it can explain why it thinks you'd like that movie, instead of just giving the recommendation.
  • Unintended ConsequencesUnintended consequences are unexpected and often negative outcomes that happen when an AI system tries to achieve its goal in a way that wasn't foreseen by its creators. The AI might be too good at its task, but in a way that causes problems. For example, if an AI is told to "maximize paperclip production" and it starts converting all available resources, including human homes, into paperclips, that's an unintended consequence.

Careers

Roles this can lead toward

AI Safety ResearcherResponsible AI EngineerAI EthicistAI Governance SpecialistAI Alignment EngineerAI Red Team LeadAI Interpretability Scientist

Student research

Published policy papers by One Young India delegates — every delegate leaves published under their own name.

← Explore the living map