Accepted at NeurIPS 2026
Attention that adapts its geometry.
DAMCHA learns an input-dependent Mahalanobis-like metric across attention heads, combining cross-head interactions with stack-wise parameter sharing.
Data-Adaptive Mahalanobis Metric Learning for Cross-Head Attention in Transformers
Get started · Explore the method · View the code · Cite the paper
Three core ideas
| Component | Role |
|---|---|
| Input-adaptive metric | A mean-pooled context conditions the metric generator. |
| Cross-head interaction | Each effective head metric sums one row of cross-head blocks. |
| Stack-wise sharing | One generator supplies metrics to the Transformer stack. |
Code and experiments
The repository contains decoder-only DAMCHA models, a Vision Transformer, five attention baselines, masked image reconstruction, conditional text generation, evaluation and profiling tools. The documentation maps these components to the paper and describes their executable configurations.
Selected paper results
Values below are reported in Table 1 of the manuscript.
| Dataset | Metric | MHA | DAMCHA |
|---|---|---|---|
| CIFAR-10 | FID@Top-50 ↓ | 91.4320 | 63.0295 |
| CIFAR-100 | FID@Top-50 ↓ | 59.6860 | 53.7560 |
| WMT | BARTScore ↑ | −8.4740 | −5.4451 |
| CommonGen | BARTScore ↑ | −5.9250 | −5.6142 |
See Experiments for task definitions, model settings and evaluation conventions.