Interpretability
AI interpretability is the study of understanding how and why an AI system produces particular behaviors or decisions.
Modern neural networks can contain billions of learned parameters, making their internal reasoning difficult to inspect directly. Interpretability research attempts to uncover useful representations, mechanisms and patterns inside these systems.
← Back to Index