MultiBLiMP v2 - Language Overview

Subject-Verb agreement prediction through decision trees. Click a language row (or the Results button) to open its create_pairs results below. The small chevron after N Keep (or the Diagnostics button) is separate: it opens just the bucket breakdown, inline, right where those items dropped out of the pipeline.
Overview

Languages shown in green have categorical agreement throughout — all samples share the label Yes.

The following languages were omitted for having too few or uninformative agreement labels (label distribution shown): Afrikaans (374), Albanian (25), Ancient Greek (1507), Ancient Hebrew (513), Armenian (1027), Aromanian (3), Assamese (8), Assyrian (1), Azerbaijani (7), Bambara (29), Basque (3893), Belarusian (4), Bengali (9), Breton (165), Bulgarian (929), Catalan (4795), Chintang (3), Classical Armenian (253), Croatian (3439), Czech (9181), Danish (1189), Dutch (2949), Egyptian (110), Esperanto (7), Estonian (2631), Faroese (12), Finnish (3496), French (4017), Galician (51), Georgian (229), German (18923), Gheg (203), Gothic (77), Greek (29), Hausa (6), Hittite (1), Icelandic (1156), Italian (5808), Karelian (2), Kazakh (31), Kiche (8), Komi Zyrian (1), Latin (2593), Latvian (2583), Ligurian (99), Lithuanian (310), Livvi (3), Low Saxon (254), Macedonian (1), Malayalam (4), Middle French (963), Moksha (5), Naija (60), Neapolitan (2), Nepali (unk: 2; No: 1), Northern Kurdish (4), Norwegian Bokmaal (2616), Norwegian Nynorsk (1860), Odia (4), Old Church Slavonic (770), Old East Slavic (1205), Old English (1), Old French (2753), Old Georgian (23), Old Occitan (654), Pashto (44), Phrygian (10), Pomak (794), Portuguese (3815), Romanian (2888), Russian (116), Sanskrit (1144), Serbian (2255), Sicilian (31), Sinhala (29), Skolt Sami (19), Slovak (3737), Slovenian (4317), Spanish (4058), Swedish (39), Tamil (185), Umbrian (1), Upper Sorbian (27), Uyghur (39), Veps (3), Western Armenian (614), Xibe (127), Zazaki (2).

Keep threshold: entropy < 0.12
Coverage 0%100%
Item loss 0%100%
Language
Distribution
Base Entropy
Reduced Entropy
Δ Entropy
DT Acc%
N RAW
N KEEP
N PAIRS