Leaderboard

This leaderboard will stay live until 31.12.2026
After that the results will remain available at

Find the task page

Find the metric implementation

Find the gold dataset

Find the competitor submissions

Track №1: Document-Level Translation with Explicit Dictionary

scored on the proper mode translations

systemchrF++ (doc)
[enpl, eseu]
chrF++ (para)
[enpl, eseu]
Term Success ?
[enpl, eseu]
Term Success (lemmatized) ?
[enpl, eseu]
COMET-22
[enpl, eseu]
CometKiwi-22
[enpl, eseu]
XCOMET-XXL
[enpl, eseu]
MetricX-24 XL ↓
[enpl, eseu]
MetricX-24 XL QE ↓
[enpl, eseu]
gold
100.0
[100.0, 100.0]
100.0
[100.0, 100.0]
96.5%
[94.2%, 98.7%]
96.4%
[93.6%, 99.1%]
96.4
[96.3, 96.5]
78.4
[81.5, 75.4]
93.1
[96.5, 89.8]
2.04
[2.26, 1.82]
3.45
[3.43, 3.47]
TARL
75.3
[77.1, 73.5]
70.4
[73.3, 67.5]
90.2%
[90.4%, 89.9%]
95.7%
[95.4%, 96.0%]
88.3
[89.0, 87.6]
77.2
[79.2, 75.1]
83.6
[85.8, 81.4]
3.92
[4.31, 3.53]
4.16
[4.49, 3.84]
COZYflash
74.4
[76.5, 72.2]
69.8
[72.9, 66.7]
88.7%
[88.8%, 88.7%]
95.4%
[94.4%, 96.3%]
88.0
[88.6, 87.4]
76.8
[79.3, 74.4]
83.1
[85.7, 80.4]
4.04
[4.33, 3.75]
4.33
[4.54, 4.12]
HW-TSC
73.8
[75.6, 72.1]
68.8
[71.9, 65.6]
91.2%
[93.0%, 89.4%]
95.3%
[95.8%, 94.8%]
87.3
[88.0, 86.7]
76.2
[78.6, 73.8]
82.2
[85.1, 79.3]
4.37
[4.75, 3.98]
4.63
[4.88, 4.39]
COZY
75.6
[75.9, 75.3]
70.9
[71.8, 70.1]
89.3%
[89.1%, 89.5%]
95.0%
[93.6%, 96.4%]
89.2
[89.6, 88.8]
77.4
[79.7, 75.0]
85.5
[87.9, 83.2]
3.58
[3.93, 3.24]
3.95
[4.13, 3.77]
TaT
73.7
[73.6, 73.8]
69.4
[69.9, 68.8]
89.9%
[96.7%, 83.0%]
94.5%
[97.2%, 91.7%]
86.8
[86.3, 87.3]
76.0
[77.1, 74.9]
80.7
[80.2, 81.2]
4.43
[5.19, 3.68]
4.66
[5.26, 4.06]
SalamandraTA
74.6
[75.9, 73.3]
69.8
[72.1, 67.5]
87.2%
[88.5%, 86.0%]
94.2%
[93.2%, 95.2%]
87.8
[88.6, 87.0]
77.0
[79.4, 74.6]
83.0
[85.7, 80.2]
4.05
[4.45, 3.66]
4.30
[4.58, 4.01]
CUNI-UFAL
74.4
[74.2, 74.5]
69.5
[69.8, 69.1]
87.7%
[87.9%, 87.4%]
93.4%
[91.6%, 95.2%]
89.1
[89.6, 88.7]
77.3
[79.5, 75.1]
85.9
[88.4, 83.5]
3.55
[3.93, 3.17]
3.88
[4.08, 3.68]
stra-mt-v3
72.0
[75.9, 68.0]
66.5
[72.0, 61.1]
84.5%
[90.0%, 79.0%]
93.1%
[94.9%, 91.3%]
△
[△, △]
△
[△, △]
△
[△, △]
△
[△, △]
△
[△, △]
TQM
70.0
[73.8, 66.1]
63.7
[69.1, 58.3]
83.9%
[87.3%, 80.6%]
89.6%
[89.8%, 89.4%]
86.5
[89.5, 83.4]
76.0
[80.0, 71.9]
80.3
[88.4, 72.3]
4.50
[3.85, 5.15]
4.49
[4.00, 4.99]
STRA-MT
69.3
[72.9, 65.8]
63.4
[68.6, 58.3]
80.7%
[86.4%, 75.1%]
89.0%
[91.7%, 86.3%]
84.1
[87.0, 81.3]
74.0
[78.3, 69.8]
74.8
[82.4, 67.3]
5.62
[5.10, 6.14]
5.44
[5.03, 5.84]
AGENTEAK
[-, 68.8]
[-, 61.3]
[-, 78.4%]
[-, 88.7%]
[-, 84.6]
[-, 72.5]
[-, 75.7]
[-, 4.71]
[-, 4.67]
Cohere-CAT+
71.1
[74.9, 67.3]
65.2
[70.7, 59.7]
80.4%
[83.8%, 77.0%]
88.0%
[84.9%, 91.1%]
87.1
[89.2, 84.9]
76.2
[79.6, 72.9]
80.6
[87.2, 74.1]
4.45
[4.14, 4.75]
4.46
[4.21, 4.71]
VNFusion
73.5
[72.8, 74.1]
67.7
[67.4, 68.1]
83.5%
[82.9%, 84.0%]
87.6%
[84.6%, 90.7%]
88.0
[88.9, 87.1]
77.6
[79.3, 75.8]
84.0
[86.7, 81.3]
3.78
[4.11, 3.44]
3.81
[4.09, 3.54]
MASC-Dec
[-, 72.7]
[-, 67.1]
[-, 82.1%]
[-, 86.8%]
[-, 84.5]
[-, 72.2]
[-, 75.2]
[-, 4.61]
[-, 4.71]
UniOR-ZSL+Term
68.8
[73.4, 64.1]
61.4
[69.4, 53.4]
81.2%
[85.5%, 76.9%]
86.4%
[87.9%, 84.9%]
83.9
[87.9, 79.9]
74.4
[79.1, 69.6]
75.8
[84.0, 67.6]
5.60
[4.72, 6.48]
5.44
[4.75, 6.14]
UniOR-RulePE
68.9
[73.5, 64.2]
60.3
[69.5, 51.0]
79.7%
[85.6%, 73.8%]
84.7%
[88.4%, 81.0%]
82.6
[87.9, 77.3]
73.2
[79.1, 67.2]
74.2
[84.1, 64.3]
5.95
[4.70, 7.21]
5.82
[4.74, 6.89]
UniOR-GenPE
68.8
[73.4, 64.2]
60.0
[69.5, 50.5]
79.4%
[85.5%, 73.2%]
84.1%
[87.7%, 80.5%]
82.5
[88.0, 77.0]
73.1
[79.2, 66.9]
74.0
[84.2, 63.7]
6.00
[4.67, 7.33]
5.85
[4.70, 7.00]
TO-fuzzy
[75.8, -]
[71.8, -]
[77.8%, -]
[80.8%, -]
[88.1, -]
[79.6, -]
[84.6, -]
[4.54, -]
[4.59, -]
TO-fuzzystem
[75.6, -]
[71.5, -]
[75.0%, -]
[78.0%, -]
[88.0, -]
[79.6, -]
[84.4, -]
[4.58, -]
[4.60, -]
UniOR-Mixed
70.4
[74.5, 66.3]
55.3
[57.8, 52.7]
71.4%
[68.3%, 74.6%]
76.4%
[70.8%, 81.9%]
76.7
[74.6, 78.8]
67.1
[65.8, 68.3]
65.7
[64.5, 67.0]
7.79
[8.97, 6.61]
7.42
[8.42, 6.42]
GRIAL-TA
[71.0, -]
[60.7, -]
[63.3%, -]
[57.3%, -]
[76.9, -]
[68.2, -]
[64.9, -]
[8.39, -]
[7.63, -]
STaR-MT
[68.6, -]
[64.0, -]
[42.4%, -]
[44.9%, -]
[85.2, -]
[78.5, -]
[78.5, -]
[5.71, -]
[5.25, -]

official competition systems systems added after competition gold = the reference translations, not a participant
△ awaiting metric evaluation on an asynchronous worker; - = direction not submitted
Scores average over the language directions in brackets; a value appears once every direction is scored.

Track №2: Document-Level Translation with Sample Bitexts

scored on the sample mode translations

systemchrF++ (doc)
[zhen, enpl, eseu]
chrF++ (para)
[zhen, enpl, eseu]
Term Success ?
[zhen, enpl, eseu]
Term Success (lemmatized) ?
[zhen, enpl, eseu]
COMET-22
[zhen, enpl, eseu]
CometKiwi-22
[zhen, enpl, eseu]
XCOMET-XXL
[zhen, enpl, eseu]
MetricX-24 XL ↓
[zhen, enpl, eseu]
MetricX-24 XL QE ↓
[zhen, enpl, eseu]
gold
100.0
[100.0, 100.0, 100.0]
100.0
[100.0, 100.0, 100.0]
98.6%
[100.0%, 96.6%, 99.3%]
93.4%
[86.3%, 94.9%, 99.1%]
95.2
[92.7, 96.3, 96.5]
78.0
[79.3, 80.8, 74.0]
91.4
[89.9, 96.4, 88.0]
1.88
[1.35, 2.37, 1.92]
3.36
[2.59, 3.65, 3.85]
CUNI-UFAL
78.1
[76.1, 81.3, 76.9]
74.3
[76.1, 78.6, 68.2]
88.5%
[95.4%, 88.5%, 81.6%]
89.2%
[88.7%, 90.4%, 88.5%]
88.7
[87.2, 91.8, 87.0]
77.4
[80.0, 80.2, 72.1]
86.0
[86.5, 91.6, 79.9]
3.22
[2.56, 3.33, 3.78]
3.64
[2.61, 3.77, 4.54]
CUNI-UFALdict-LLM
74.9
[74.5, 75.0, 75.2]
71.3
[74.5, 71.1, 68.2]
87.0%
[94.5%, 83.2%, 83.4%]
89.0%
[88.3%, 87.2%, 91.5%]
88.5
[86.9, 90.0, 88.5]
77.8
[80.2, 79.5, 73.8]
85.9
[86.3, 89.5, 81.9]
3.25
[2.62, 3.76, 3.37]
3.51
[2.55, 3.93, 4.05]
COZY
77.5
[75.7, 79.5, 77.3]
74.2
[75.7, 76.5, 70.4]
88.3%
[95.4%, 86.1%, 83.3%]
88.9%
[88.8%, 87.5%, 90.5%]
88.9
[87.1, 90.9, 88.7]
78.0
[80.1, 80.3, 73.7]
86.4
[86.7, 90.8, 81.7]
3.18
[2.60, 3.55, 3.38]
3.51
[2.63, 3.81, 4.08]
stra-mt-v3
[-, 75.8, 68.5]
[-, 72.2, 59.2]
[-, 83.0%, 55.9%]
[-, 87.7%, 69.1%]
△
[-, △, △]
△
[-, △, △]
△
[-, △, △]
△
[-, △, △]
△
[-, △, △]
COZYflash
76.8
[74.5, 80.7, 75.1]
73.4
[74.5, 77.9, 67.8]
86.5%
[93.7%, 86.1%, 79.8%]
87.2%
[87.1%, 87.8%, 86.7%]
88.3
[86.8, 90.6, 87.6]
77.7
[80.0, 79.8, 73.3]
84.9
[86.0, 89.0, 79.8]
3.37
[2.68, 3.76, 3.68]
3.66
[2.66, 4.06, 4.27]
CUNI-UFALE2E
78.1
[76.3, 81.1, 76.9]
74.7
[76.3, 78.3, 69.4]
88.1%
[95.8%, 87.0%, 81.5%]
87.1%
[89.0%, 83.9%, 88.3%]
88.9
[87.1, 91.6, 88.0]
77.9
[80.2, 80.5, 73.0]
86.5
[86.5, 91.8, 81.2]
3.14
[2.59, 3.33, 3.51]
3.54
[2.61, 3.73, 4.28]
HW-TSC
74.0
[73.1, 76.4, 72.4]
70.2
[73.1, 73.0, 64.5]
78.8%
[91.7%, 76.8%, 68.0%]
85.1%
[84.7%, 88.1%, 82.4%]
86.5
[86.0, 87.4, 86.1]
76.3
[79.3, 77.4, 72.3]
82.0
[84.1, 85.2, 76.8]
4.10
[3.02, 4.93, 4.34]
4.24
[2.92, 5.01, 4.78]
TaT
75.1
[74.6, 75.1, 75.7]
71.6
[74.6, 71.5, 68.6]
83.3%
[92.8%, 78.6%, 78.3%]
84.0%
[88.4%, 80.6%, 83.1%]
87.6
[86.4, 88.2, 88.3]
77.7
[80.1, 78.9, 74.1]
84.0
[86.0, 85.3, 80.8]
3.58
[2.73, 4.50, 3.52]
3.75
[2.62, 4.50, 4.14]
TARL
73.7
[72.6, 75.6, 72.9]
69.7
[72.6, 71.8, 64.8]
76.1%
[92.2%, 67.4%, 68.8%]
80.4%
[85.5%, 77.2%, 78.6%]
85.8
[84.9, 88.0, 84.6]
76.5
[78.9, 78.7, 71.9]
81.6
[83.0, 85.5, 76.2]
4.15
[3.34, 4.68, 4.44]
4.16
[3.06, 4.64, 4.77]
TO-G4terms
[-, 75.4, -]
[-, 71.5, -]
[-, 77.4%, -]
[-, 79.8%, -]
[-, 88.5, -]
[-, 79.9, -]
[-, 86.2, -]
[-, 4.33, -]
[-, 4.34, -]
TO-TMtop15
[-, 80.1, -]
[-, 77.2, -]
[-, 80.4%, -]
[-, 78.7%, -]
[-, 89.7, -]
[-, 79.3, -]
[-, 89.0, -]
[-, 4.00, -]
[-, 4.35, -]
Cohere-CAT+
73.4
[71.9, 78.0, 70.4]
69.0
[71.9, 74.6, 60.4]
75.7%
[89.4%, 76.0%, 61.6%]
75.2%
[82.4%, 73.4%, 69.9%]
86.8
[85.4, 90.4, 84.7]
77.2
[79.8, 80.4, 71.5]
82.7
[84.8, 90.2, 73.2]
3.78
[2.92, 3.65, 4.76]
3.75
[2.62, 3.77, 4.87]
TQM
73.0
[71.8, 78.8, 68.5]
68.4
[71.8, 75.7, 57.6]
76.2%
[90.9%, 82.1%, 55.7%]
74.2%
[80.7%, 79.2%, 62.7%]
86.1
[85.0, 90.2, 83.2]
76.7
[79.1, 79.5, 71.5]
80.4
[81.7, 88.5, 71.0]
4.16
[3.23, 3.95, 5.31]
4.06
[2.93, 4.19, 5.06]
TO-TMtop5
[-, 78.6, -]
[-, 75.4, -]
[-, 73.2%, -]
[-, 73.7%, -]
[-, 89.9, -]
[-, 80.2, -]
[-, 88.3, -]
[-, 4.01, -]
[-, 4.17, -]
VNFusion
73.8
[70.1, 75.1, 76.3]
69.8
[70.1, 70.8, 68.6]
75.5%
[86.2%, 70.3%, 70.0%]
73.4%
[78.0%, 66.4%, 75.8%]
87.0
[84.6, 88.9, 87.4]
77.9
[79.1, 79.9, 74.8]
83.3
[82.5, 87.3, 80.2]
3.61
[3.23, 4.08, 3.52]
3.47
[2.80, 3.89, 3.72]
UniVie-HAITrans
[-, 73.0, 67.5]
[-, 68.4, 55.3]
[-, 68.3%, 46.4%]
[-, 62.1%, 53.5%]
[-, 85.3, 78.9]
[-, 75.4, 67.3]
[-, 77.4, 59.4]
[-, 5.71, 7.46]
[-, 5.44, 6.82]
STaR-MT
67.3
[70.1, 68.1, 63.6]
62.7
[70.1, 63.6, 54.4]
52.5%
[86.5%, 40.6%, 30.5%]
53.1%
[80.6%, 42.8%, 35.9%]
83.9
[84.2, 85.5, 82.0]
76.5
[79.2, 78.8, 71.4]
77.1
[82.0, 79.8, 69.5]
4.87
[3.31, 5.49, 5.81]
4.33
[2.83, 4.84, 5.32]
STRA-MT
65.9
[65.5, 67.3, 64.8]
60.9
[65.5, 62.4, 54.8]
45.4%
[67.7%, 36.7%, 31.7%]
48.0%
[64.9%, 38.7%, 40.3%]
80.9
[80.2, 84.0, 78.6]
73.7
[75.8, 77.7, 67.6]
68.7
[68.2, 76.3, 61.5]
6.19
[5.04, 6.21, 7.31]
5.34
[3.94, 5.41, 6.69]
GRIAL-TA
[-, 62.2, -]
[-, 50.2, -]
[-, 33.4%, -]
[-, 32.1%, -]
[-, 64.4, -]
[-, 55.5, -]
[-, 43.8, -]
[-, 13.13, -]
[-, 11.92, -]

official competition systems systems added after competition gold = the reference translations, not a participant
△ awaiting metric evaluation on an asynchronous worker; - = direction not submitted
Scores average over the language directions in brackets; a value appears once every direction is scored.