MultiBLiMP v2 - Language Overview

Subject-Verb agreement prediction through decision trees. Click a language row (or the Results button) to open its create_pairs results below. The small chevron after N Keep (or the Diagnostics button) is separate: it opens just the bucket breakdown, inline, right where those items dropped out of the pipeline.
Overview

Languages shown in green have categorical agreement throughout — all samples share the label Yes.

The following languages were omitted for having too few or uninformative agreement labels (label distribution shown): Abaza (69), Afrikaans (1805), Alemannic (884), Bambara (1763), Basque (2298), Bavarian (816), Bokota (237), Cantonese (598), Cebuano (127), Chinese (15147), Classical Chinese (39438), Coptic (1140), Danish (5776), Esperanto (185), Gorontalo (1), Guarani (1), Gujarati (107), Gwichin (27), Haitian Creole (6403), Ika (170), Indonesian (7251), Japanese (6473), Javanese (784), Kadiweu (43), Khoekhoe (2205), Korean (23796), Luxembourgish (20), Makurap (10), Malayalam (126), Maltese (1626), Middle French (5989), Munduruku (62), Naga (355), Nenets (33), Northwest Gbaya (367), Norwegian Bokmaal (17120), Norwegian Nynorsk (16336), Occitan (921), Old French (15398), Old Irish (10), Old Turkish (2), Paumari (23), Pesh (83), Phrygian (115), Sanskrit (9099), Shanghainese (670), Sinhala (unk: 81; No: 1), South Levantine Arabic (23), Tagalog (107), Telugu (922), Thai (5729), Tswana (20), Umbrian (27), Vietnamese (3326), Warlpiri (56), Western Sierra Puebla Nahuatl (1061), Xibe (561), Yiddish (2187), Yoruba (721), Yupik (163), Zaar (336).

Keep threshold: entropy < 0.12
Coverage 0%100%
Item loss 0%100%
Language
Distribution
Base Entropy
Reduced Entropy
Δ Entropy
DT Acc%
N RAW
N KEEP
N PAIRS