Our Mission

Human understanding should grow with machine capability.

The International Consortium for Interpretable AI will develop the science needed for humans to understand, and thus effectively harness, machine intelligence.

By building a solid foundation of evidence-based scientific tools to clarify the internal mechanisms of AI, the Consortium will enable people to control AI behavior and learn from its discoveries.

The imperative

Artificial intelligence has inverted the relationship between people and computers. Through the history of computing, programming has been humanity’s way of shaping and understanding automation. But today’s neural networks are trained, not programmed. Their internal mechanisms emerge from fitting vast amounts of data, leaving us with the challenge of reverse-engineering the learned programs that underlie their behavior.

Because human understanding is no longer a prerequisite to automation, computer science faces a new imperative: we must ensure that as AI capabilities rise, human agency grows.

The urgency of interpretability is stark: AI systems are already being adopted across society. AI has begun to displace human work across every field touched by software, mathematics, and writing. But ceding human thinking to machine intelligence will make it difficult for people to maintain real responsibility.

Government, science, engineering, and medicine are all predicated on the assumption that complex decision-making intertwines accountability with insight. As AI continues to take on intellectual work, humanity must find ways to incisively understand how AI reaches its conclusions. We must dispel our ignorance and uphold our intentions, so we can take responsibility for results.

A scientific foundation

The science of AI interpretability is nascent. While many studies have observed the structure of AI computations through a variety of lenses, the field lacks consensus around shared definitions, methods, and standards of evidence.

ICINA will build a solid scientific foundation for empirical, evidence-driven AI interpretability by convening researchers, developing infrastructure, and creating scientific commons. The work will be motivated by the urgent mission of increasing human insight and responsibility as AI becomes more autonomous.

Three scientific goals

01

Auditing AI

Can we build a lie detector for AI?

In both routine and high-stakes use, there is a gap between what an AI system says and what it knows internally. Closing this gap is a fundamental challenge for AI safety.

The Consortium will develop rigorous methods for inspecting AI computations to detect deception, hidden goals, censorship, hallucination, reward hacking, and other serious failures missed by ordinary evaluation. It will serve as an impartial shepherd for benchmarks and experimental protocols, and as a training ground for expertise in monitoring and mitigating deception.

02

Understanding intelligence

What is thinking?

With AI, we now have silicon brains whose every computation can be observed, altered, and repeated. We face a historic opportunity to investigate ancient questions about the nature of reasoning and thought.

The Consortium will dissect the structure of neural representations and trace causal pathways underlying cognitive mechanisms. We will develop measurable definitions for cognitive phenomena and evidence-based methods to understand the origin of the profound capabilities of large neural networks.

03

Empowering people

How can AI amplify human agency?

The most valuable aspect of AI is its knowledge beyond human knowledge, but the science of AI will have failed if it only serves to make AI systems superhuman. People cannot answer for decisions they do not understand.

The Consortium will develop methods for identifying and explaining AI discoveries at the edge of human knowledge, translating AI knowledge into human insights, and measuring whether those insights improve human judgment.

Our principles

The Consortium is driven by an obligation to ask important questions even when unprofitable, to hold fast to the optimism that hard problems can be solved, and to celebrate the serious play that connects understanding with action.

For humanity to be resilient in the face of unexpected AI challenges, we must translate insight into the technical means to correct AI, and build a culture in which people expect to exercise that power. The practice of interrogating and changing AI systems must be practical.

AI will not become understandable to us simply because we created it. Bringing AI within reach of human understanding will require us to build a new science.