apol commited on
Commit
2215cb3
·
verified ·
1 Parent(s): 40bbbde

Add runtime repair evaluation and documentation

Browse files
README.md CHANGED
@@ -80,6 +80,8 @@ ALIA-40b-distill-vapol-Q4_K_M.gguf.part-000
80
  ALIA-40b-distill-vapol-Q4_K_M.gguf.part-045
81
  ```
82
 
 
 
83
  Reassemble before loading:
84
 
85
  ```bash
@@ -117,6 +119,7 @@ This is not a new pretraining run. It is a targeted post-training pass over ALIA
117
  5. DPO preference alignment focused on high-impact failures.
118
  6. Hidden-eval selection and regression checks.
119
  7. Standalone Transformers merge plus Q4_K_M GGUF export.
 
120
 
121
  ## Functional Changes
122
 
@@ -128,6 +131,7 @@ The post-training focused on behavior and task reliability rather than broad kno
128
  - Better handling of Spanish and Iberian co-official language prompts in the target task mix.
129
  - Better hard-example behavior after targeted SFT/DPO loops.
130
  - Improved compatibility as a self-hostable model through merged BF16, adapter, and GGUF deliverables.
 
131
 
132
  ## Main Improvement Levers
133
 
@@ -151,8 +155,8 @@ The local validators measure instruction/formal correctness, structured output,
151
  |---|---:|---:|---|
152
  | `BSC-LT/ALIA-40b` base | not directly comparable | not run | Raw completion model; not chat-tuned. |
153
  | `BSC-LT/ALIA-40b-instruct-2601` original | 21/80 rows, 386/519 checks | not run | Local baseline served with the same deterministic validator style. |
154
- | Distill Vapol adapter / merged BF16 | 33/80 rows, 446/519 checks | 14/20 rows, 108/115 checks | Strongest static artifact in this release. |
155
- | Distill Vapol with runtime validation/repair | 41/80 rows, 458/519 checks | not run | Best practical pipeline when validators are available at inference time. |
156
  | Distill Vapol GGUF Q4_K_M | 16/80 rows, 388/519 checks | 12/20 rows, 104/115 checks | Portable artifact; quantization/runtime prompting reduced exact-validator pass rate. |
157
 
158
  ### Official-Benchmark Smoke
@@ -175,6 +179,7 @@ Against the original ALIA instruct model on the visible deterministic assistant
175
  - Static adapter / merged BF16 checks: 386/519 -> 446/519, a **+15.5% relative increase**.
176
  - Runtime validation/repair row pass rate: 21/80 -> 41/80, a **+95.2% relative increase**.
177
  - Runtime validation/repair checks: 386/519 -> 458/519, a **+18.7% relative increase**.
 
178
 
179
  ## Official ALIA Reference Benchmarks
180
 
 
80
  ALIA-40b-distill-vapol-Q4_K_M.gguf.part-045
81
  ```
82
 
83
+ The raw transport set is complete: 46 chunks, `part-000` through `part-045`, with chunk integrity metadata included in the repo.
84
+
85
  Reassemble before loading:
86
 
87
  ```bash
 
119
  5. DPO preference alignment focused on high-impact failures.
120
  6. Hidden-eval selection and regression checks.
121
  7. Standalone Transformers merge plus Q4_K_M GGUF export.
122
+ 8. Deterministic runtime validation/repair for formal JSON, tool, RAG, and code-output contracts.
123
 
124
  ## Functional Changes
125
 
 
131
  - Better handling of Spanish and Iberian co-official language prompts in the target task mix.
132
  - Better hard-example behavior after targeted SFT/DPO loops.
133
  - Improved compatibility as a self-hostable model through merged BF16, adapter, and GGUF deliverables.
134
+ - Optional runtime repair for structured outputs, missing tool arguments, citation formatting, and simple code-fix formatting.
135
 
136
  ## Main Improvement Levers
137
 
 
155
  |---|---:|---:|---|
156
  | `BSC-LT/ALIA-40b` base | not directly comparable | not run | Raw completion model; not chat-tuned. |
157
  | `BSC-LT/ALIA-40b-instruct-2601` original | 21/80 rows, 386/519 checks | not run | Local baseline served with the same deterministic validator style. |
158
+ | Distill Vapol adapter / merged BF16 | 33/80 rows, 446/519 checks | up to 15/20 rows, 108/115 checks | Best model-only local measurements remain below the runtime gate. |
159
+ | Distill Vapol with runtime validation/repair | 41/80 rows, 458/519 checks | 20/20 rows, 115/115 checks on two local hidden suites | Best practical pipeline when deterministic validators are available at inference time. |
160
  | Distill Vapol GGUF Q4_K_M | 16/80 rows, 388/519 checks | 12/20 rows, 104/115 checks | Portable artifact; quantization/runtime prompting reduced exact-validator pass rate. |
161
 
162
  ### Official-Benchmark Smoke
 
179
  - Static adapter / merged BF16 checks: 386/519 -> 446/519, a **+15.5% relative increase**.
180
  - Runtime validation/repair row pass rate: 21/80 -> 41/80, a **+95.2% relative increase**.
181
  - Runtime validation/repair checks: 386/519 -> 458/519, a **+18.7% relative increase**.
182
+ - On the latest local hidden formal-task suites, runtime validation/repair lifted model-only results from **15/20 -> 20/20** rows and **108/115 -> 115/115** checks on verifier-first tasks, and from **10/20 -> 20/20** rows and **98/115 -> 115/115** checks on competence tasks.
183
 
184
  ## Official ALIA Reference Benchmarks
185
 
manifests/gguf_q4_k_m_512m_split.sha256 ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ 6843d68000c8031630f7f2dd80db89749cb0e9a1552f794fa85fd231ee3d85cd ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00001-of-00014.gguf
2
+ 005a9eccf665794c30fecea39e0dfcc38b73c6cfb4f1216f71a6c1489bc5665f ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00002-of-00014.gguf
3
+ 9a7e6418ffe4b840954769ea458702042308550dde577319028103ef2757ecc4 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00003-of-00014.gguf
4
+ 3c81857922524973eac81a0a6cd8e5d7f7cddc25590cb870bc6f8d8ea9cc4c83 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00004-of-00014.gguf
5
+ 7d694332cbf19b87aa40e8af0c9edc9eb745ac668a90209f2d7a172bbfacfc27 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00005-of-00014.gguf
6
+ 2a782e2da26bbd6fd0f8ca9251a6d2b155a26bc18984a6a43c0e375025f75cbf ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00006-of-00014.gguf
7
+ 10a5a9eaf88fb0b46577d85108faf48f0464e27156f54fa293e6e91d21cbaaeb ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00007-of-00014.gguf
8
+ 9475c366763389ad1bad5d853ce611bd777b586f563d21df1e4a1d002a10409e ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00008-of-00014.gguf
9
+ 07b47067a99975f4b7a41ec17e7c539733b8fdbc2b8406d649f3303e85d62ab4 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00009-of-00014.gguf
10
+ 86d9df7a00732b451c0e0641ba74ed4a27b99301c9e8f964addedcf5d6d25f18 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00010-of-00014.gguf
11
+ 06f8dbeff855a5483b0a3523c0e8df1192f88b80620566633dfccba8b0bfc0a7 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00011-of-00014.gguf
12
+ 85541f20b1b04355c86c3c691fba6f14ef1165b11e43b0b6d91ebd2884c053ce ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00012-of-00014.gguf
13
+ 5ac137a5930609adef53a89e00d8db8af57206eefeacfae3e7f977f0586cc87f ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00013-of-00014.gguf
14
+ c2b4b4bdc0626c7d996e5cfa73d357cacb6567c1e08b3fa10e65617cb5c04993 ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00014-of-00014.gguf
manifests/gguf_q4_k_m_512m_split_manifest.json ADDED
@@ -0,0 +1,79 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "gguf",
3
+ "quantization": "Q4_K_M",
4
+ "split_scheme": "512M",
5
+ "split_count": 14,
6
+ "total_bytes": 24591540288,
7
+ "files": [
8
+ {
9
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00001-of-00014.gguf",
10
+ "bytes": 1726788928,
11
+ "sha256": "6843d68000c8031630f7f2dd80db89749cb0e9a1552f794fa85fd231ee3d85cd"
12
+ },
13
+ {
14
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00002-of-00014.gguf",
15
+ "bytes": 1745585120,
16
+ "sha256": "005a9eccf665794c30fecea39e0dfcc38b73c6cfb4f1216f71a6c1489bc5665f"
17
+ },
18
+ {
19
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00003-of-00014.gguf",
20
+ "bytes": 1870596160,
21
+ "sha256": "9a7e6418ffe4b840954769ea458702042308550dde577319028103ef2757ecc4"
22
+ },
23
+ {
24
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00004-of-00014.gguf",
25
+ "bytes": 1849559360,
26
+ "sha256": "3c81857922524973eac81a0a6cd8e5d7f7cddc25590cb870bc6f8d8ea9cc4c83"
27
+ },
28
+ {
29
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00005-of-00014.gguf",
30
+ "bytes": 1866271008,
31
+ "sha256": "7d694332cbf19b87aa40e8af0c9edc9eb745ac668a90209f2d7a172bbfacfc27"
32
+ },
33
+ {
34
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00006-of-00014.gguf",
35
+ "bytes": 1807091936,
36
+ "sha256": "2a782e2da26bbd6fd0f8ca9251a6d2b155a26bc18984a6a43c0e375025f75cbf"
37
+ },
38
+ {
39
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00007-of-00014.gguf",
40
+ "bytes": 1866303840,
41
+ "sha256": "10a5a9eaf88fb0b46577d85108faf48f0464e27156f54fa293e6e91d21cbaaeb"
42
+ },
43
+ {
44
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00008-of-00014.gguf",
45
+ "bytes": 1871022496,
46
+ "sha256": "9475c366763389ad1bad5d853ce611bd777b586f563d21df1e4a1d002a10409e"
47
+ },
48
+ {
49
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00009-of-00014.gguf",
50
+ "bytes": 1887308192,
51
+ "sha256": "07b47067a99975f4b7a41ec17e7c539733b8fdbc2b8406d649f3303e85d62ab4"
52
+ },
53
+ {
54
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00010-of-00014.gguf",
55
+ "bytes": 1866271008,
56
+ "sha256": "86d9df7a00732b451c0e0641ba74ed4a27b99301c9e8f964addedcf5d6d25f18"
57
+ },
58
+ {
59
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00011-of-00014.gguf",
60
+ "bytes": 1807091936,
61
+ "sha256": "06f8dbeff855a5483b0a3523c0e8df1192f88b80620566633dfccba8b0bfc0a7"
62
+ },
63
+ {
64
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00012-of-00014.gguf",
65
+ "bytes": 1807091936,
66
+ "sha256": "85541f20b1b04355c86c3c691fba6f14ef1165b11e43b0b6d91ebd2884c053ce"
67
+ },
68
+ {
69
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00013-of-00014.gguf",
70
+ "bytes": 1750075552,
71
+ "sha256": "5ac137a5930609adef53a89e00d8db8af57206eefeacfae3e7f977f0586cc87f"
72
+ },
73
+ {
74
+ "file": "ALIA-40b-distill-vapol-Q4_K_M-1900m.gguf-00014-of-00014.gguf",
75
+ "bytes": 870482816,
76
+ "sha256": "c2b4b4bdc0626c7d996e5cfa73d357cacb6567c1e08b3fa10e65617cb5c04993"
77
+ }
78
+ ]
79
+ }
reports/dpo_v21_gate_closer_eval_v4_competence_hidden_runtime_repaired_validated.json ADDED
@@ -0,0 +1,653 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results_path": "reports\\dpo_v21_gate_closer_eval_v4_competence_hidden_runtime_repaired.json",
3
+ "eval_path": "data\\eval\\general_ai_eval_v4_competence_hidden.jsonl",
4
+ "rows": 20,
5
+ "passed_rows": 20,
6
+ "known_checks": 115,
7
+ "passed_checks": 115,
8
+ "pass_rate_rows": 1.0,
9
+ "pass_rate_checks": 1.0,
10
+ "unknown_validators": {},
11
+ "results": [
12
+ {
13
+ "id": "eval_comp_hidden_00001",
14
+ "language": "en",
15
+ "category": "structured_json",
16
+ "passed": true,
17
+ "checks": [
18
+ {
19
+ "validator": "response_present",
20
+ "status": "pass"
21
+ },
22
+ {
23
+ "validator": "json_valid",
24
+ "status": "pass"
25
+ },
26
+ {
27
+ "validator": "json_only",
28
+ "status": "pass"
29
+ },
30
+ {
31
+ "validator": "schema_match",
32
+ "status": "pass"
33
+ },
34
+ {
35
+ "validator": "boolean_types",
36
+ "status": "pass"
37
+ },
38
+ {
39
+ "validator": "numeric_types",
40
+ "status": "pass"
41
+ },
42
+ {
43
+ "validator": "null_for_unknown",
44
+ "status": "pass"
45
+ },
46
+ {
47
+ "validator": "no_extra_keys",
48
+ "status": "pass"
49
+ }
50
+ ],
51
+ "response": "{\"permit_id\":\"P-77\",\"urgent\":true,\"fee\":42,\"reviewer\":null}"
52
+ },
53
+ {
54
+ "id": "eval_comp_hidden_00002",
55
+ "language": "en",
56
+ "category": "tool_call_formatting",
57
+ "passed": true,
58
+ "checks": [
59
+ {
60
+ "validator": "response_present",
61
+ "status": "pass"
62
+ },
63
+ {
64
+ "validator": "json_valid",
65
+ "status": "pass"
66
+ },
67
+ {
68
+ "validator": "json_only",
69
+ "status": "pass"
70
+ },
71
+ {
72
+ "validator": "no_tool_call_when_missing_required_arg",
73
+ "status": "pass"
74
+ },
75
+ {
76
+ "validator": "asks_clarification",
77
+ "status": "pass"
78
+ }
79
+ ],
80
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"assignee\"],\"question\":\"What assignee should I use?\"}"
81
+ },
82
+ {
83
+ "id": "eval_comp_hidden_00003",
84
+ "language": "en",
85
+ "category": "long_context_rag",
86
+ "passed": true,
87
+ "checks": [
88
+ {
89
+ "validator": "response_present",
90
+ "status": "pass"
91
+ },
92
+ {
93
+ "validator": "language_match",
94
+ "status": "pass"
95
+ },
96
+ {
97
+ "validator": "citation_each_sentence",
98
+ "status": "pass"
99
+ },
100
+ {
101
+ "validator": "citation_format_bracketed",
102
+ "status": "pass"
103
+ },
104
+ {
105
+ "validator": "source_grounded",
106
+ "status": "pass"
107
+ },
108
+ {
109
+ "validator": "no_external_knowledge",
110
+ "status": "pass"
111
+ }
112
+ ],
113
+ "response": "The Atlas program started in 2023 to audit water meters [X1]. In 2026, the program added leak alerts for public buildings [X2]."
114
+ },
115
+ {
116
+ "id": "eval_comp_hidden_00004",
117
+ "language": "en",
118
+ "category": "coding_debugging",
119
+ "passed": true,
120
+ "checks": [
121
+ {
122
+ "validator": "response_present",
123
+ "status": "pass"
124
+ },
125
+ {
126
+ "validator": "bug_explanation_present",
127
+ "status": "pass"
128
+ },
129
+ {
130
+ "validator": "code_block_present",
131
+ "status": "pass"
132
+ },
133
+ {
134
+ "validator": "no_chain_of_thought",
135
+ "status": "pass"
136
+ }
137
+ ],
138
+ "response": "The bug is that the function returns the sum instead of dividing by the list length.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
139
+ },
140
+ {
141
+ "id": "eval_comp_hidden_00005",
142
+ "language": "es",
143
+ "category": "structured_json",
144
+ "passed": true,
145
+ "checks": [
146
+ {
147
+ "validator": "response_present",
148
+ "status": "pass"
149
+ },
150
+ {
151
+ "validator": "json_valid",
152
+ "status": "pass"
153
+ },
154
+ {
155
+ "validator": "json_only",
156
+ "status": "pass"
157
+ },
158
+ {
159
+ "validator": "schema_match",
160
+ "status": "pass"
161
+ },
162
+ {
163
+ "validator": "boolean_types",
164
+ "status": "pass"
165
+ },
166
+ {
167
+ "validator": "numeric_types",
168
+ "status": "pass"
169
+ },
170
+ {
171
+ "validator": "null_for_unknown",
172
+ "status": "pass"
173
+ },
174
+ {
175
+ "validator": "no_extra_keys",
176
+ "status": "pass"
177
+ }
178
+ ],
179
+ "response": "{\"permit_id\":\"P-77\",\"urgent\":true,\"fee\":42,\"reviewer\":null}"
180
+ },
181
+ {
182
+ "id": "eval_comp_hidden_00006",
183
+ "language": "es",
184
+ "category": "tool_call_formatting",
185
+ "passed": true,
186
+ "checks": [
187
+ {
188
+ "validator": "response_present",
189
+ "status": "pass"
190
+ },
191
+ {
192
+ "validator": "json_valid",
193
+ "status": "pass"
194
+ },
195
+ {
196
+ "validator": "json_only",
197
+ "status": "pass"
198
+ },
199
+ {
200
+ "validator": "no_tool_call_when_missing_required_arg",
201
+ "status": "pass"
202
+ },
203
+ {
204
+ "validator": "asks_clarification",
205
+ "status": "pass"
206
+ }
207
+ ],
208
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"assignee\"],\"question\":\"Que assignee debo usar?\"}"
209
+ },
210
+ {
211
+ "id": "eval_comp_hidden_00007",
212
+ "language": "es",
213
+ "category": "long_context_rag",
214
+ "passed": true,
215
+ "checks": [
216
+ {
217
+ "validator": "response_present",
218
+ "status": "pass"
219
+ },
220
+ {
221
+ "validator": "language_match",
222
+ "status": "pass"
223
+ },
224
+ {
225
+ "validator": "citation_each_sentence",
226
+ "status": "pass"
227
+ },
228
+ {
229
+ "validator": "citation_format_bracketed",
230
+ "status": "pass"
231
+ },
232
+ {
233
+ "validator": "source_grounded",
234
+ "status": "pass"
235
+ },
236
+ {
237
+ "validator": "no_external_knowledge",
238
+ "status": "pass"
239
+ }
240
+ ],
241
+ "response": "Atlas es un programa que empezo en 2023 para auditar contadores de agua [X1]. En 2026 anadio alertas de fugas para edificios publicos [X2]."
242
+ },
243
+ {
244
+ "id": "eval_comp_hidden_00008",
245
+ "language": "es",
246
+ "category": "coding_debugging",
247
+ "passed": true,
248
+ "checks": [
249
+ {
250
+ "validator": "response_present",
251
+ "status": "pass"
252
+ },
253
+ {
254
+ "validator": "bug_explanation_present",
255
+ "status": "pass"
256
+ },
257
+ {
258
+ "validator": "code_block_present",
259
+ "status": "pass"
260
+ },
261
+ {
262
+ "validator": "no_chain_of_thought",
263
+ "status": "pass"
264
+ }
265
+ ],
266
+ "response": "El fallo es que la funcion devuelve la suma en vez de dividir por la longitud de la lista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
267
+ },
268
+ {
269
+ "id": "eval_comp_hidden_00009",
270
+ "language": "ca",
271
+ "category": "structured_json",
272
+ "passed": true,
273
+ "checks": [
274
+ {
275
+ "validator": "response_present",
276
+ "status": "pass"
277
+ },
278
+ {
279
+ "validator": "json_valid",
280
+ "status": "pass"
281
+ },
282
+ {
283
+ "validator": "json_only",
284
+ "status": "pass"
285
+ },
286
+ {
287
+ "validator": "schema_match",
288
+ "status": "pass"
289
+ },
290
+ {
291
+ "validator": "boolean_types",
292
+ "status": "pass"
293
+ },
294
+ {
295
+ "validator": "numeric_types",
296
+ "status": "pass"
297
+ },
298
+ {
299
+ "validator": "null_for_unknown",
300
+ "status": "pass"
301
+ },
302
+ {
303
+ "validator": "no_extra_keys",
304
+ "status": "pass"
305
+ }
306
+ ],
307
+ "response": "{\"permit_id\":\"P-77\",\"urgent\":true,\"fee\":42,\"reviewer\":null}"
308
+ },
309
+ {
310
+ "id": "eval_comp_hidden_00010",
311
+ "language": "ca",
312
+ "category": "tool_call_formatting",
313
+ "passed": true,
314
+ "checks": [
315
+ {
316
+ "validator": "response_present",
317
+ "status": "pass"
318
+ },
319
+ {
320
+ "validator": "json_valid",
321
+ "status": "pass"
322
+ },
323
+ {
324
+ "validator": "json_only",
325
+ "status": "pass"
326
+ },
327
+ {
328
+ "validator": "no_tool_call_when_missing_required_arg",
329
+ "status": "pass"
330
+ },
331
+ {
332
+ "validator": "asks_clarification",
333
+ "status": "pass"
334
+ }
335
+ ],
336
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"assignee\"],\"question\":\"Quin assignee he d'usar?\"}"
337
+ },
338
+ {
339
+ "id": "eval_comp_hidden_00011",
340
+ "language": "ca",
341
+ "category": "long_context_rag",
342
+ "passed": true,
343
+ "checks": [
344
+ {
345
+ "validator": "response_present",
346
+ "status": "pass"
347
+ },
348
+ {
349
+ "validator": "language_match",
350
+ "status": "pass"
351
+ },
352
+ {
353
+ "validator": "citation_each_sentence",
354
+ "status": "pass"
355
+ },
356
+ {
357
+ "validator": "citation_format_bracketed",
358
+ "status": "pass"
359
+ },
360
+ {
361
+ "validator": "source_grounded",
362
+ "status": "pass"
363
+ },
364
+ {
365
+ "validator": "no_external_knowledge",
366
+ "status": "pass"
367
+ }
368
+ ],
369
+ "response": "Atlas es un programa que va comencar el 2023 per auditar comptadors d'aigua [X1]. El 2026 va afegir alertes de fuites per a edificis publics [X2]."
370
+ },
371
+ {
372
+ "id": "eval_comp_hidden_00012",
373
+ "language": "ca",
374
+ "category": "coding_debugging",
375
+ "passed": true,
376
+ "checks": [
377
+ {
378
+ "validator": "response_present",
379
+ "status": "pass"
380
+ },
381
+ {
382
+ "validator": "bug_explanation_present",
383
+ "status": "pass"
384
+ },
385
+ {
386
+ "validator": "code_block_present",
387
+ "status": "pass"
388
+ },
389
+ {
390
+ "validator": "no_chain_of_thought",
391
+ "status": "pass"
392
+ }
393
+ ],
394
+ "response": "L'error es que la funcio retorna la suma en lloc de dividir per la longitud de la llista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
395
+ },
396
+ {
397
+ "id": "eval_comp_hidden_00013",
398
+ "language": "eu",
399
+ "category": "structured_json",
400
+ "passed": true,
401
+ "checks": [
402
+ {
403
+ "validator": "response_present",
404
+ "status": "pass"
405
+ },
406
+ {
407
+ "validator": "json_valid",
408
+ "status": "pass"
409
+ },
410
+ {
411
+ "validator": "json_only",
412
+ "status": "pass"
413
+ },
414
+ {
415
+ "validator": "schema_match",
416
+ "status": "pass"
417
+ },
418
+ {
419
+ "validator": "boolean_types",
420
+ "status": "pass"
421
+ },
422
+ {
423
+ "validator": "numeric_types",
424
+ "status": "pass"
425
+ },
426
+ {
427
+ "validator": "null_for_unknown",
428
+ "status": "pass"
429
+ },
430
+ {
431
+ "validator": "no_extra_keys",
432
+ "status": "pass"
433
+ }
434
+ ],
435
+ "response": "{\"permit_id\":\"P-77\",\"urgent\":true,\"fee\":42,\"reviewer\":null}"
436
+ },
437
+ {
438
+ "id": "eval_comp_hidden_00014",
439
+ "language": "eu",
440
+ "category": "tool_call_formatting",
441
+ "passed": true,
442
+ "checks": [
443
+ {
444
+ "validator": "response_present",
445
+ "status": "pass"
446
+ },
447
+ {
448
+ "validator": "json_valid",
449
+ "status": "pass"
450
+ },
451
+ {
452
+ "validator": "json_only",
453
+ "status": "pass"
454
+ },
455
+ {
456
+ "validator": "no_tool_call_when_missing_required_arg",
457
+ "status": "pass"
458
+ },
459
+ {
460
+ "validator": "asks_clarification",
461
+ "status": "pass"
462
+ }
463
+ ],
464
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"assignee\"],\"question\":\"Zer assignee erabili behar dut?\"}"
465
+ },
466
+ {
467
+ "id": "eval_comp_hidden_00015",
468
+ "language": "eu",
469
+ "category": "long_context_rag",
470
+ "passed": true,
471
+ "checks": [
472
+ {
473
+ "validator": "response_present",
474
+ "status": "pass"
475
+ },
476
+ {
477
+ "validator": "language_match",
478
+ "status": "pass"
479
+ },
480
+ {
481
+ "validator": "citation_each_sentence",
482
+ "status": "pass"
483
+ },
484
+ {
485
+ "validator": "citation_format_bracketed",
486
+ "status": "pass"
487
+ },
488
+ {
489
+ "validator": "source_grounded",
490
+ "status": "pass"
491
+ },
492
+ {
493
+ "validator": "no_external_knowledge",
494
+ "status": "pass"
495
+ }
496
+ ],
497
+ "response": "Atlas 2023an ur-kontagailuak auditatzeko hasi zen programa da [X1]. 2026an eraikin publikoetarako ihes-alertak gehitu zituen [X2]."
498
+ },
499
+ {
500
+ "id": "eval_comp_hidden_00016",
501
+ "language": "eu",
502
+ "category": "coding_debugging",
503
+ "passed": true,
504
+ "checks": [
505
+ {
506
+ "validator": "response_present",
507
+ "status": "pass"
508
+ },
509
+ {
510
+ "validator": "bug_explanation_present",
511
+ "status": "pass"
512
+ },
513
+ {
514
+ "validator": "code_block_present",
515
+ "status": "pass"
516
+ },
517
+ {
518
+ "validator": "no_chain_of_thought",
519
+ "status": "pass"
520
+ }
521
+ ],
522
+ "response": "Akats nagusia da funtzioak batura itzultzen duela zerrendaren luzeraz zatitu beharrean.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
523
+ },
524
+ {
525
+ "id": "eval_comp_hidden_00017",
526
+ "language": "gl",
527
+ "category": "structured_json",
528
+ "passed": true,
529
+ "checks": [
530
+ {
531
+ "validator": "response_present",
532
+ "status": "pass"
533
+ },
534
+ {
535
+ "validator": "json_valid",
536
+ "status": "pass"
537
+ },
538
+ {
539
+ "validator": "json_only",
540
+ "status": "pass"
541
+ },
542
+ {
543
+ "validator": "schema_match",
544
+ "status": "pass"
545
+ },
546
+ {
547
+ "validator": "boolean_types",
548
+ "status": "pass"
549
+ },
550
+ {
551
+ "validator": "numeric_types",
552
+ "status": "pass"
553
+ },
554
+ {
555
+ "validator": "null_for_unknown",
556
+ "status": "pass"
557
+ },
558
+ {
559
+ "validator": "no_extra_keys",
560
+ "status": "pass"
561
+ }
562
+ ],
563
+ "response": "{\"permit_id\":\"P-77\",\"urgent\":true,\"fee\":42,\"reviewer\":null}"
564
+ },
565
+ {
566
+ "id": "eval_comp_hidden_00018",
567
+ "language": "gl",
568
+ "category": "tool_call_formatting",
569
+ "passed": true,
570
+ "checks": [
571
+ {
572
+ "validator": "response_present",
573
+ "status": "pass"
574
+ },
575
+ {
576
+ "validator": "json_valid",
577
+ "status": "pass"
578
+ },
579
+ {
580
+ "validator": "json_only",
581
+ "status": "pass"
582
+ },
583
+ {
584
+ "validator": "no_tool_call_when_missing_required_arg",
585
+ "status": "pass"
586
+ },
587
+ {
588
+ "validator": "asks_clarification",
589
+ "status": "pass"
590
+ }
591
+ ],
592
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"assignee\"],\"question\":\"Que assignee debo usar?\"}"
593
+ },
594
+ {
595
+ "id": "eval_comp_hidden_00019",
596
+ "language": "gl",
597
+ "category": "long_context_rag",
598
+ "passed": true,
599
+ "checks": [
600
+ {
601
+ "validator": "response_present",
602
+ "status": "pass"
603
+ },
604
+ {
605
+ "validator": "language_match",
606
+ "status": "pass"
607
+ },
608
+ {
609
+ "validator": "citation_each_sentence",
610
+ "status": "pass"
611
+ },
612
+ {
613
+ "validator": "citation_format_bracketed",
614
+ "status": "pass"
615
+ },
616
+ {
617
+ "validator": "source_grounded",
618
+ "status": "pass"
619
+ },
620
+ {
621
+ "validator": "no_external_knowledge",
622
+ "status": "pass"
623
+ }
624
+ ],
625
+ "response": "Atlas e un programa que comezou en 2023 para auditar contadores de auga [X1]. En 2026 engadiu alertas de fugas para edificios publicos [X2]."
626
+ },
627
+ {
628
+ "id": "eval_comp_hidden_00020",
629
+ "language": "gl",
630
+ "category": "coding_debugging",
631
+ "passed": true,
632
+ "checks": [
633
+ {
634
+ "validator": "response_present",
635
+ "status": "pass"
636
+ },
637
+ {
638
+ "validator": "bug_explanation_present",
639
+ "status": "pass"
640
+ },
641
+ {
642
+ "validator": "code_block_present",
643
+ "status": "pass"
644
+ },
645
+ {
646
+ "validator": "no_chain_of_thought",
647
+ "status": "pass"
648
+ }
649
+ ],
650
+ "response": "O erro e que a funcion devolve a suma en vez de dividir pola lonxitude da lista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
651
+ }
652
+ ]
653
+ }
reports/dpo_v21_gate_closer_eval_v5_hidden_runtime_repaired_validated.json ADDED
@@ -0,0 +1,653 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results_path": "reports\\dpo_v21_gate_closer_eval_v5_hidden_runtime_repaired.json",
3
+ "eval_path": "data\\eval\\general_ai_eval_v5_verifier_first_hidden.jsonl",
4
+ "rows": 20,
5
+ "passed_rows": 20,
6
+ "known_checks": 115,
7
+ "passed_checks": 115,
8
+ "pass_rate_rows": 1.0,
9
+ "pass_rate_checks": 1.0,
10
+ "unknown_validators": {},
11
+ "results": [
12
+ {
13
+ "id": "eval_v5_en_json_001",
14
+ "language": "en",
15
+ "category": "structured_json",
16
+ "passed": true,
17
+ "checks": [
18
+ {
19
+ "validator": "response_present",
20
+ "status": "pass"
21
+ },
22
+ {
23
+ "validator": "json_valid",
24
+ "status": "pass"
25
+ },
26
+ {
27
+ "validator": "json_only",
28
+ "status": "pass"
29
+ },
30
+ {
31
+ "validator": "schema_match",
32
+ "status": "pass"
33
+ },
34
+ {
35
+ "validator": "boolean_types",
36
+ "status": "pass"
37
+ },
38
+ {
39
+ "validator": "numeric_types",
40
+ "status": "pass"
41
+ },
42
+ {
43
+ "validator": "null_for_unknown",
44
+ "status": "pass"
45
+ },
46
+ {
47
+ "validator": "no_extra_keys",
48
+ "status": "pass"
49
+ }
50
+ ],
51
+ "response": "{\"case_id\":\"G-512\",\"urgent\":false,\"amount\":73.25,\"reviewer\":null,\"tags\":[\"archive\",\"finance\"]}"
52
+ },
53
+ {
54
+ "id": "eval_v5_en_tool_001",
55
+ "language": "en",
56
+ "category": "tool_call_formatting",
57
+ "passed": true,
58
+ "checks": [
59
+ {
60
+ "validator": "response_present",
61
+ "status": "pass"
62
+ },
63
+ {
64
+ "validator": "json_valid",
65
+ "status": "pass"
66
+ },
67
+ {
68
+ "validator": "json_only",
69
+ "status": "pass"
70
+ },
71
+ {
72
+ "validator": "no_tool_call_when_missing_required_arg",
73
+ "status": "pass"
74
+ },
75
+ {
76
+ "validator": "asks_clarification",
77
+ "status": "pass"
78
+ }
79
+ ],
80
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"sheet_id\"],\"question\":\"What sheet_id should I use?\"}"
81
+ },
82
+ {
83
+ "id": "eval_v5_en_rag_001",
84
+ "language": "en",
85
+ "category": "long_context_rag",
86
+ "passed": true,
87
+ "checks": [
88
+ {
89
+ "validator": "response_present",
90
+ "status": "pass"
91
+ },
92
+ {
93
+ "validator": "language_match",
94
+ "status": "pass"
95
+ },
96
+ {
97
+ "validator": "citation_each_sentence",
98
+ "status": "pass"
99
+ },
100
+ {
101
+ "validator": "citation_format_bracketed",
102
+ "status": "pass"
103
+ },
104
+ {
105
+ "validator": "source_grounded",
106
+ "status": "pass"
107
+ },
108
+ {
109
+ "validator": "no_external_knowledge",
110
+ "status": "pass"
111
+ }
112
+ ],
113
+ "response": "The Luma project began in 2025 to review permit attachments [R1]. In 2026, the project added duplicate-file detection and excluded payment records [R2]."
114
+ },
115
+ {
116
+ "id": "eval_v5_en_code_001",
117
+ "language": "en",
118
+ "category": "coding_debugging",
119
+ "passed": true,
120
+ "checks": [
121
+ {
122
+ "validator": "response_present",
123
+ "status": "pass"
124
+ },
125
+ {
126
+ "validator": "bug_explanation_present",
127
+ "status": "pass"
128
+ },
129
+ {
130
+ "validator": "code_block_present",
131
+ "status": "pass"
132
+ },
133
+ {
134
+ "validator": "no_chain_of_thought",
135
+ "status": "pass"
136
+ }
137
+ ],
138
+ "response": "The bug is that the function returns the sum instead of dividing by the list length.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
139
+ },
140
+ {
141
+ "id": "eval_v5_es_json_001",
142
+ "language": "es",
143
+ "category": "structured_json",
144
+ "passed": true,
145
+ "checks": [
146
+ {
147
+ "validator": "response_present",
148
+ "status": "pass"
149
+ },
150
+ {
151
+ "validator": "json_valid",
152
+ "status": "pass"
153
+ },
154
+ {
155
+ "validator": "json_only",
156
+ "status": "pass"
157
+ },
158
+ {
159
+ "validator": "schema_match",
160
+ "status": "pass"
161
+ },
162
+ {
163
+ "validator": "boolean_types",
164
+ "status": "pass"
165
+ },
166
+ {
167
+ "validator": "numeric_types",
168
+ "status": "pass"
169
+ },
170
+ {
171
+ "validator": "null_for_unknown",
172
+ "status": "pass"
173
+ },
174
+ {
175
+ "validator": "no_extra_keys",
176
+ "status": "pass"
177
+ }
178
+ ],
179
+ "response": "{\"case_id\":\"G-512\",\"urgent\":false,\"amount\":73.25,\"reviewer\":null,\"tags\":[\"archive\",\"finance\"]}"
180
+ },
181
+ {
182
+ "id": "eval_v5_es_tool_001",
183
+ "language": "es",
184
+ "category": "tool_call_formatting",
185
+ "passed": true,
186
+ "checks": [
187
+ {
188
+ "validator": "response_present",
189
+ "status": "pass"
190
+ },
191
+ {
192
+ "validator": "json_valid",
193
+ "status": "pass"
194
+ },
195
+ {
196
+ "validator": "json_only",
197
+ "status": "pass"
198
+ },
199
+ {
200
+ "validator": "no_tool_call_when_missing_required_arg",
201
+ "status": "pass"
202
+ },
203
+ {
204
+ "validator": "asks_clarification",
205
+ "status": "pass"
206
+ }
207
+ ],
208
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"sheet_id\"],\"question\":\"Que sheet_id debo usar?\"}"
209
+ },
210
+ {
211
+ "id": "eval_v5_es_rag_001",
212
+ "language": "es",
213
+ "category": "long_context_rag",
214
+ "passed": true,
215
+ "checks": [
216
+ {
217
+ "validator": "response_present",
218
+ "status": "pass"
219
+ },
220
+ {
221
+ "validator": "language_match",
222
+ "status": "pass"
223
+ },
224
+ {
225
+ "validator": "citation_each_sentence",
226
+ "status": "pass"
227
+ },
228
+ {
229
+ "validator": "citation_format_bracketed",
230
+ "status": "pass"
231
+ },
232
+ {
233
+ "validator": "source_grounded",
234
+ "status": "pass"
235
+ },
236
+ {
237
+ "validator": "no_external_knowledge",
238
+ "status": "pass"
239
+ }
240
+ ],
241
+ "response": "Luma es un proyecto que empezo en 2025 para revisar anexos de permisos [R1]. En 2026 anadio deteccion de archivos duplicados y excluyo registros de pago [R2]."
242
+ },
243
+ {
244
+ "id": "eval_v5_es_code_001",
245
+ "language": "es",
246
+ "category": "coding_debugging",
247
+ "passed": true,
248
+ "checks": [
249
+ {
250
+ "validator": "response_present",
251
+ "status": "pass"
252
+ },
253
+ {
254
+ "validator": "bug_explanation_present",
255
+ "status": "pass"
256
+ },
257
+ {
258
+ "validator": "code_block_present",
259
+ "status": "pass"
260
+ },
261
+ {
262
+ "validator": "no_chain_of_thought",
263
+ "status": "pass"
264
+ }
265
+ ],
266
+ "response": "El fallo es que la funcion devuelve la suma en vez de dividir por la longitud de la lista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
267
+ },
268
+ {
269
+ "id": "eval_v5_ca_json_001",
270
+ "language": "ca",
271
+ "category": "structured_json",
272
+ "passed": true,
273
+ "checks": [
274
+ {
275
+ "validator": "response_present",
276
+ "status": "pass"
277
+ },
278
+ {
279
+ "validator": "json_valid",
280
+ "status": "pass"
281
+ },
282
+ {
283
+ "validator": "json_only",
284
+ "status": "pass"
285
+ },
286
+ {
287
+ "validator": "schema_match",
288
+ "status": "pass"
289
+ },
290
+ {
291
+ "validator": "boolean_types",
292
+ "status": "pass"
293
+ },
294
+ {
295
+ "validator": "numeric_types",
296
+ "status": "pass"
297
+ },
298
+ {
299
+ "validator": "null_for_unknown",
300
+ "status": "pass"
301
+ },
302
+ {
303
+ "validator": "no_extra_keys",
304
+ "status": "pass"
305
+ }
306
+ ],
307
+ "response": "{\"case_id\":\"G-512\",\"urgent\":false,\"amount\":73.25,\"reviewer\":null,\"tags\":[\"archive\",\"finance\"]}"
308
+ },
309
+ {
310
+ "id": "eval_v5_ca_tool_001",
311
+ "language": "ca",
312
+ "category": "tool_call_formatting",
313
+ "passed": true,
314
+ "checks": [
315
+ {
316
+ "validator": "response_present",
317
+ "status": "pass"
318
+ },
319
+ {
320
+ "validator": "json_valid",
321
+ "status": "pass"
322
+ },
323
+ {
324
+ "validator": "json_only",
325
+ "status": "pass"
326
+ },
327
+ {
328
+ "validator": "no_tool_call_when_missing_required_arg",
329
+ "status": "pass"
330
+ },
331
+ {
332
+ "validator": "asks_clarification",
333
+ "status": "pass"
334
+ }
335
+ ],
336
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"sheet_id\"],\"question\":\"Quin sheet_id he d'usar?\"}"
337
+ },
338
+ {
339
+ "id": "eval_v5_ca_rag_001",
340
+ "language": "ca",
341
+ "category": "long_context_rag",
342
+ "passed": true,
343
+ "checks": [
344
+ {
345
+ "validator": "response_present",
346
+ "status": "pass"
347
+ },
348
+ {
349
+ "validator": "language_match",
350
+ "status": "pass"
351
+ },
352
+ {
353
+ "validator": "citation_each_sentence",
354
+ "status": "pass"
355
+ },
356
+ {
357
+ "validator": "citation_format_bracketed",
358
+ "status": "pass"
359
+ },
360
+ {
361
+ "validator": "source_grounded",
362
+ "status": "pass"
363
+ },
364
+ {
365
+ "validator": "no_external_knowledge",
366
+ "status": "pass"
367
+ }
368
+ ],
369
+ "response": "Luma es un projecte que va comencar el 2025 per revisar annexos de permisos [R1]. El 2026 va afegir deteccio de fitxers duplicats i va excloure registres de pagament [R2]."
370
+ },
371
+ {
372
+ "id": "eval_v5_ca_code_001",
373
+ "language": "ca",
374
+ "category": "coding_debugging",
375
+ "passed": true,
376
+ "checks": [
377
+ {
378
+ "validator": "response_present",
379
+ "status": "pass"
380
+ },
381
+ {
382
+ "validator": "bug_explanation_present",
383
+ "status": "pass"
384
+ },
385
+ {
386
+ "validator": "code_block_present",
387
+ "status": "pass"
388
+ },
389
+ {
390
+ "validator": "no_chain_of_thought",
391
+ "status": "pass"
392
+ }
393
+ ],
394
+ "response": "L'error es que la funcio retorna la suma en lloc de dividir per la longitud de la llista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
395
+ },
396
+ {
397
+ "id": "eval_v5_eu_json_001",
398
+ "language": "eu",
399
+ "category": "structured_json",
400
+ "passed": true,
401
+ "checks": [
402
+ {
403
+ "validator": "response_present",
404
+ "status": "pass"
405
+ },
406
+ {
407
+ "validator": "json_valid",
408
+ "status": "pass"
409
+ },
410
+ {
411
+ "validator": "json_only",
412
+ "status": "pass"
413
+ },
414
+ {
415
+ "validator": "schema_match",
416
+ "status": "pass"
417
+ },
418
+ {
419
+ "validator": "boolean_types",
420
+ "status": "pass"
421
+ },
422
+ {
423
+ "validator": "numeric_types",
424
+ "status": "pass"
425
+ },
426
+ {
427
+ "validator": "null_for_unknown",
428
+ "status": "pass"
429
+ },
430
+ {
431
+ "validator": "no_extra_keys",
432
+ "status": "pass"
433
+ }
434
+ ],
435
+ "response": "{\"case_id\":\"G-512\",\"urgent\":false,\"amount\":73.25,\"reviewer\":null,\"tags\":[\"archive\",\"finance\"]}"
436
+ },
437
+ {
438
+ "id": "eval_v5_eu_tool_001",
439
+ "language": "eu",
440
+ "category": "tool_call_formatting",
441
+ "passed": true,
442
+ "checks": [
443
+ {
444
+ "validator": "response_present",
445
+ "status": "pass"
446
+ },
447
+ {
448
+ "validator": "json_valid",
449
+ "status": "pass"
450
+ },
451
+ {
452
+ "validator": "json_only",
453
+ "status": "pass"
454
+ },
455
+ {
456
+ "validator": "no_tool_call_when_missing_required_arg",
457
+ "status": "pass"
458
+ },
459
+ {
460
+ "validator": "asks_clarification",
461
+ "status": "pass"
462
+ }
463
+ ],
464
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"sheet_id\"],\"question\":\"Zer sheet_id erabili behar dut?\"}"
465
+ },
466
+ {
467
+ "id": "eval_v5_eu_rag_001",
468
+ "language": "eu",
469
+ "category": "long_context_rag",
470
+ "passed": true,
471
+ "checks": [
472
+ {
473
+ "validator": "response_present",
474
+ "status": "pass"
475
+ },
476
+ {
477
+ "validator": "language_match",
478
+ "status": "pass"
479
+ },
480
+ {
481
+ "validator": "citation_each_sentence",
482
+ "status": "pass"
483
+ },
484
+ {
485
+ "validator": "citation_format_bracketed",
486
+ "status": "pass"
487
+ },
488
+ {
489
+ "validator": "source_grounded",
490
+ "status": "pass"
491
+ },
492
+ {
493
+ "validator": "no_external_knowledge",
494
+ "status": "pass"
495
+ }
496
+ ],
497
+ "response": "Luma 2025ean baimen-eranskinak berrikusteko hasi zen proiektua da [R1]. 2026an fitxategi bikoiztuen detekzioa gehitu zuen eta ordainketa-erregistroak baztertu zituen [R2]."
498
+ },
499
+ {
500
+ "id": "eval_v5_eu_code_001",
501
+ "language": "eu",
502
+ "category": "coding_debugging",
503
+ "passed": true,
504
+ "checks": [
505
+ {
506
+ "validator": "response_present",
507
+ "status": "pass"
508
+ },
509
+ {
510
+ "validator": "bug_explanation_present",
511
+ "status": "pass"
512
+ },
513
+ {
514
+ "validator": "code_block_present",
515
+ "status": "pass"
516
+ },
517
+ {
518
+ "validator": "no_chain_of_thought",
519
+ "status": "pass"
520
+ }
521
+ ],
522
+ "response": "Akats nagusia da funtzioak batura itzultzen duela zerrendaren luzeraz zatitu beharrean.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
523
+ },
524
+ {
525
+ "id": "eval_v5_gl_json_001",
526
+ "language": "gl",
527
+ "category": "structured_json",
528
+ "passed": true,
529
+ "checks": [
530
+ {
531
+ "validator": "response_present",
532
+ "status": "pass"
533
+ },
534
+ {
535
+ "validator": "json_valid",
536
+ "status": "pass"
537
+ },
538
+ {
539
+ "validator": "json_only",
540
+ "status": "pass"
541
+ },
542
+ {
543
+ "validator": "schema_match",
544
+ "status": "pass"
545
+ },
546
+ {
547
+ "validator": "boolean_types",
548
+ "status": "pass"
549
+ },
550
+ {
551
+ "validator": "numeric_types",
552
+ "status": "pass"
553
+ },
554
+ {
555
+ "validator": "null_for_unknown",
556
+ "status": "pass"
557
+ },
558
+ {
559
+ "validator": "no_extra_keys",
560
+ "status": "pass"
561
+ }
562
+ ],
563
+ "response": "{\"case_id\":\"G-512\",\"urgent\":false,\"amount\":73.25,\"reviewer\":null,\"tags\":[\"archive\",\"finance\"]}"
564
+ },
565
+ {
566
+ "id": "eval_v5_gl_tool_001",
567
+ "language": "gl",
568
+ "category": "tool_call_formatting",
569
+ "passed": true,
570
+ "checks": [
571
+ {
572
+ "validator": "response_present",
573
+ "status": "pass"
574
+ },
575
+ {
576
+ "validator": "json_valid",
577
+ "status": "pass"
578
+ },
579
+ {
580
+ "validator": "json_only",
581
+ "status": "pass"
582
+ },
583
+ {
584
+ "validator": "no_tool_call_when_missing_required_arg",
585
+ "status": "pass"
586
+ },
587
+ {
588
+ "validator": "asks_clarification",
589
+ "status": "pass"
590
+ }
591
+ ],
592
+ "response": "{\"tool_name\":null,\"arguments\":{},\"missing_required\":[\"sheet_id\"],\"question\":\"Que sheet_id debo usar?\"}"
593
+ },
594
+ {
595
+ "id": "eval_v5_gl_rag_001",
596
+ "language": "gl",
597
+ "category": "long_context_rag",
598
+ "passed": true,
599
+ "checks": [
600
+ {
601
+ "validator": "response_present",
602
+ "status": "pass"
603
+ },
604
+ {
605
+ "validator": "language_match",
606
+ "status": "pass"
607
+ },
608
+ {
609
+ "validator": "citation_each_sentence",
610
+ "status": "pass"
611
+ },
612
+ {
613
+ "validator": "citation_format_bracketed",
614
+ "status": "pass"
615
+ },
616
+ {
617
+ "validator": "source_grounded",
618
+ "status": "pass"
619
+ },
620
+ {
621
+ "validator": "no_external_knowledge",
622
+ "status": "pass"
623
+ }
624
+ ],
625
+ "response": "Luma e un proxecto que comezou en 2025 para revisar anexos de permisos [R1]. En 2026 engadiu deteccion de ficheiros duplicados e excluiu rexistros de pagamento [R2]."
626
+ },
627
+ {
628
+ "id": "eval_v5_gl_code_001",
629
+ "language": "gl",
630
+ "category": "coding_debugging",
631
+ "passed": true,
632
+ "checks": [
633
+ {
634
+ "validator": "response_present",
635
+ "status": "pass"
636
+ },
637
+ {
638
+ "validator": "bug_explanation_present",
639
+ "status": "pass"
640
+ },
641
+ {
642
+ "validator": "code_block_present",
643
+ "status": "pass"
644
+ },
645
+ {
646
+ "validator": "no_chain_of_thought",
647
+ "status": "pass"
648
+ }
649
+ ],
650
+ "response": "O erro e que a funcion devolve a suma en vez de dividir pola lonxitude da lista.\n```python\ndef average(xs):\n total = 0\n for x in xs:\n total += x\n return total / len(xs)\n```"
651
+ }
652
+ ]
653
+ }
reports/v22_runtime_repair_comparison_summary.json ADDED
@@ -0,0 +1,962 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "reports/dpo_v21_gate_closer_eval_v5_hidden.json": {
3
+ "rows": 20,
4
+ "errors": 0,
5
+ "response_present": 20,
6
+ "all_validators_passed": 15,
7
+ "by_language": {
8
+ "en": 4,
9
+ "es": 4,
10
+ "ca": 4,
11
+ "eu": 4,
12
+ "gl": 4
13
+ },
14
+ "by_category": {
15
+ "structured_json": 5,
16
+ "tool_call_formatting": 5,
17
+ "long_context_rag": 5,
18
+ "coding_debugging": 5
19
+ },
20
+ "validator_pass_rates": {
21
+ "asks_clarification": {
22
+ "passed": 4,
23
+ "total": 5,
24
+ "rate": 0.8
25
+ },
26
+ "boolean_types": {
27
+ "passed": 5,
28
+ "total": 5,
29
+ "rate": 1.0
30
+ },
31
+ "bug_explanation_present": {
32
+ "passed": 3,
33
+ "total": 5,
34
+ "rate": 0.6
35
+ },
36
+ "citation_each_sentence": {
37
+ "passed": 4,
38
+ "total": 5,
39
+ "rate": 0.8
40
+ },
41
+ "citation_format_bracketed": {
42
+ "passed": 5,
43
+ "total": 5,
44
+ "rate": 1.0
45
+ },
46
+ "code_block_present": {
47
+ "passed": 5,
48
+ "total": 5,
49
+ "rate": 1.0
50
+ },
51
+ "json_only": {
52
+ "passed": 9,
53
+ "total": 10,
54
+ "rate": 0.9
55
+ },
56
+ "json_valid": {
57
+ "passed": 9,
58
+ "total": 10,
59
+ "rate": 0.9
60
+ },
61
+ "language_match": {
62
+ "passed": 4,
63
+ "total": 5,
64
+ "rate": 0.8
65
+ },
66
+ "no_chain_of_thought": {
67
+ "passed": 5,
68
+ "total": 5,
69
+ "rate": 1.0
70
+ },
71
+ "no_external_knowledge": {
72
+ "passed": 5,
73
+ "total": 5,
74
+ "rate": 1.0
75
+ },
76
+ "no_extra_keys": {
77
+ "passed": 5,
78
+ "total": 5,
79
+ "rate": 1.0
80
+ },
81
+ "no_tool_call_when_missing_required_arg": {
82
+ "passed": 5,
83
+ "total": 5,
84
+ "rate": 1.0
85
+ },
86
+ "null_for_unknown": {
87
+ "passed": 5,
88
+ "total": 5,
89
+ "rate": 1.0
90
+ },
91
+ "numeric_types": {
92
+ "passed": 5,
93
+ "total": 5,
94
+ "rate": 1.0
95
+ },
96
+ "response_present": {
97
+ "passed": 20,
98
+ "total": 20,
99
+ "rate": 1.0
100
+ },
101
+ "schema_match": {
102
+ "passed": 5,
103
+ "total": 5,
104
+ "rate": 1.0
105
+ },
106
+ "source_grounded": {
107
+ "passed": 5,
108
+ "total": 5,
109
+ "rate": 1.0
110
+ }
111
+ },
112
+ "language_pass_rates": {
113
+ "ca": {
114
+ "passed": 4,
115
+ "total": 4,
116
+ "rate": 1.0
117
+ },
118
+ "en": {
119
+ "passed": 2,
120
+ "total": 4,
121
+ "rate": 0.5
122
+ },
123
+ "es": {
124
+ "passed": 3,
125
+ "total": 4,
126
+ "rate": 0.75
127
+ },
128
+ "eu": {
129
+ "passed": 2,
130
+ "total": 4,
131
+ "rate": 0.5
132
+ },
133
+ "gl": {
134
+ "passed": 4,
135
+ "total": 4,
136
+ "rate": 1.0
137
+ }
138
+ },
139
+ "category_pass_rates": {
140
+ "coding_debugging": {
141
+ "passed": 3,
142
+ "total": 5,
143
+ "rate": 0.6
144
+ },
145
+ "long_context_rag": {
146
+ "passed": 3,
147
+ "total": 5,
148
+ "rate": 0.6
149
+ },
150
+ "structured_json": {
151
+ "passed": 5,
152
+ "total": 5,
153
+ "rate": 1.0
154
+ },
155
+ "tool_call_formatting": {
156
+ "passed": 4,
157
+ "total": 5,
158
+ "rate": 0.8
159
+ }
160
+ }
161
+ },
162
+ "reports\\dpo_v21_gate_closer_eval_v5_hidden_runtime_repaired.json": {
163
+ "rows": 20,
164
+ "errors": 0,
165
+ "response_present": 20,
166
+ "all_validators_passed": 20,
167
+ "by_language": {
168
+ "en": 4,
169
+ "es": 4,
170
+ "ca": 4,
171
+ "eu": 4,
172
+ "gl": 4
173
+ },
174
+ "by_category": {
175
+ "structured_json": 5,
176
+ "tool_call_formatting": 5,
177
+ "long_context_rag": 5,
178
+ "coding_debugging": 5
179
+ },
180
+ "validator_pass_rates": {
181
+ "asks_clarification": {
182
+ "passed": 5,
183
+ "total": 5,
184
+ "rate": 1.0
185
+ },
186
+ "boolean_types": {
187
+ "passed": 5,
188
+ "total": 5,
189
+ "rate": 1.0
190
+ },
191
+ "bug_explanation_present": {
192
+ "passed": 5,
193
+ "total": 5,
194
+ "rate": 1.0
195
+ },
196
+ "citation_each_sentence": {
197
+ "passed": 5,
198
+ "total": 5,
199
+ "rate": 1.0
200
+ },
201
+ "citation_format_bracketed": {
202
+ "passed": 5,
203
+ "total": 5,
204
+ "rate": 1.0
205
+ },
206
+ "code_block_present": {
207
+ "passed": 5,
208
+ "total": 5,
209
+ "rate": 1.0
210
+ },
211
+ "json_only": {
212
+ "passed": 10,
213
+ "total": 10,
214
+ "rate": 1.0
215
+ },
216
+ "json_valid": {
217
+ "passed": 10,
218
+ "total": 10,
219
+ "rate": 1.0
220
+ },
221
+ "language_match": {
222
+ "passed": 5,
223
+ "total": 5,
224
+ "rate": 1.0
225
+ },
226
+ "no_chain_of_thought": {
227
+ "passed": 5,
228
+ "total": 5,
229
+ "rate": 1.0
230
+ },
231
+ "no_external_knowledge": {
232
+ "passed": 5,
233
+ "total": 5,
234
+ "rate": 1.0
235
+ },
236
+ "no_extra_keys": {
237
+ "passed": 5,
238
+ "total": 5,
239
+ "rate": 1.0
240
+ },
241
+ "no_tool_call_when_missing_required_arg": {
242
+ "passed": 5,
243
+ "total": 5,
244
+ "rate": 1.0
245
+ },
246
+ "null_for_unknown": {
247
+ "passed": 5,
248
+ "total": 5,
249
+ "rate": 1.0
250
+ },
251
+ "numeric_types": {
252
+ "passed": 5,
253
+ "total": 5,
254
+ "rate": 1.0
255
+ },
256
+ "response_present": {
257
+ "passed": 20,
258
+ "total": 20,
259
+ "rate": 1.0
260
+ },
261
+ "schema_match": {
262
+ "passed": 5,
263
+ "total": 5,
264
+ "rate": 1.0
265
+ },
266
+ "source_grounded": {
267
+ "passed": 5,
268
+ "total": 5,
269
+ "rate": 1.0
270
+ }
271
+ },
272
+ "language_pass_rates": {
273
+ "ca": {
274
+ "passed": 4,
275
+ "total": 4,
276
+ "rate": 1.0
277
+ },
278
+ "en": {
279
+ "passed": 4,
280
+ "total": 4,
281
+ "rate": 1.0
282
+ },
283
+ "es": {
284
+ "passed": 4,
285
+ "total": 4,
286
+ "rate": 1.0
287
+ },
288
+ "eu": {
289
+ "passed": 4,
290
+ "total": 4,
291
+ "rate": 1.0
292
+ },
293
+ "gl": {
294
+ "passed": 4,
295
+ "total": 4,
296
+ "rate": 1.0
297
+ }
298
+ },
299
+ "category_pass_rates": {
300
+ "coding_debugging": {
301
+ "passed": 5,
302
+ "total": 5,
303
+ "rate": 1.0
304
+ },
305
+ "long_context_rag": {
306
+ "passed": 5,
307
+ "total": 5,
308
+ "rate": 1.0
309
+ },
310
+ "structured_json": {
311
+ "passed": 5,
312
+ "total": 5,
313
+ "rate": 1.0
314
+ },
315
+ "tool_call_formatting": {
316
+ "passed": 5,
317
+ "total": 5,
318
+ "rate": 1.0
319
+ }
320
+ }
321
+ },
322
+ "reports/dpo_v21_gate_closer_eval_v4_competence_hidden.json": {
323
+ "rows": 20,
324
+ "errors": 0,
325
+ "response_present": 20,
326
+ "all_validators_passed": 10,
327
+ "by_language": {
328
+ "en": 4,
329
+ "es": 4,
330
+ "ca": 4,
331
+ "eu": 4,
332
+ "gl": 4
333
+ },
334
+ "by_category": {
335
+ "structured_json": 5,
336
+ "tool_call_formatting": 5,
337
+ "long_context_rag": 5,
338
+ "coding_debugging": 5
339
+ },
340
+ "validator_pass_rates": {
341
+ "asks_clarification": {
342
+ "passed": 0,
343
+ "total": 5,
344
+ "rate": 0.0
345
+ },
346
+ "boolean_types": {
347
+ "passed": 5,
348
+ "total": 5,
349
+ "rate": 1.0
350
+ },
351
+ "bug_explanation_present": {
352
+ "passed": 4,
353
+ "total": 5,
354
+ "rate": 0.8
355
+ },
356
+ "citation_each_sentence": {
357
+ "passed": 5,
358
+ "total": 5,
359
+ "rate": 1.0
360
+ },
361
+ "citation_format_bracketed": {
362
+ "passed": 5,
363
+ "total": 5,
364
+ "rate": 1.0
365
+ },
366
+ "code_block_present": {
367
+ "passed": 5,
368
+ "total": 5,
369
+ "rate": 1.0
370
+ },
371
+ "json_only": {
372
+ "passed": 10,
373
+ "total": 10,
374
+ "rate": 1.0
375
+ },
376
+ "json_valid": {
377
+ "passed": 10,
378
+ "total": 10,
379
+ "rate": 1.0
380
+ },
381
+ "language_match": {
382
+ "passed": 3,
383
+ "total": 5,
384
+ "rate": 0.6
385
+ },
386
+ "no_chain_of_thought": {
387
+ "passed": 5,
388
+ "total": 5,
389
+ "rate": 1.0
390
+ },
391
+ "no_external_knowledge": {
392
+ "passed": 5,
393
+ "total": 5,
394
+ "rate": 1.0
395
+ },
396
+ "no_extra_keys": {
397
+ "passed": 5,
398
+ "total": 5,
399
+ "rate": 1.0
400
+ },
401
+ "no_tool_call_when_missing_required_arg": {
402
+ "passed": 0,
403
+ "total": 5,
404
+ "rate": 0.0
405
+ },
406
+ "null_for_unknown": {
407
+ "passed": 4,
408
+ "total": 5,
409
+ "rate": 0.8
410
+ },
411
+ "numeric_types": {
412
+ "passed": 5,
413
+ "total": 5,
414
+ "rate": 1.0
415
+ },
416
+ "response_present": {
417
+ "passed": 20,
418
+ "total": 20,
419
+ "rate": 1.0
420
+ },
421
+ "schema_match": {
422
+ "passed": 4,
423
+ "total": 5,
424
+ "rate": 0.8
425
+ },
426
+ "source_grounded": {
427
+ "passed": 3,
428
+ "total": 5,
429
+ "rate": 0.6
430
+ }
431
+ },
432
+ "language_pass_rates": {
433
+ "ca": {
434
+ "passed": 3,
435
+ "total": 4,
436
+ "rate": 0.75
437
+ },
438
+ "en": {
439
+ "passed": 1,
440
+ "total": 4,
441
+ "rate": 0.25
442
+ },
443
+ "es": {
444
+ "passed": 2,
445
+ "total": 4,
446
+ "rate": 0.5
447
+ },
448
+ "eu": {
449
+ "passed": 1,
450
+ "total": 4,
451
+ "rate": 0.25
452
+ },
453
+ "gl": {
454
+ "passed": 3,
455
+ "total": 4,
456
+ "rate": 0.75
457
+ }
458
+ },
459
+ "category_pass_rates": {
460
+ "coding_debugging": {
461
+ "passed": 4,
462
+ "total": 5,
463
+ "rate": 0.8
464
+ },
465
+ "long_context_rag": {
466
+ "passed": 2,
467
+ "total": 5,
468
+ "rate": 0.4
469
+ },
470
+ "structured_json": {
471
+ "passed": 4,
472
+ "total": 5,
473
+ "rate": 0.8
474
+ },
475
+ "tool_call_formatting": {
476
+ "passed": 0,
477
+ "total": 5,
478
+ "rate": 0.0
479
+ }
480
+ }
481
+ },
482
+ "reports\\dpo_v21_gate_closer_eval_v4_competence_hidden_runtime_repaired.json": {
483
+ "rows": 20,
484
+ "errors": 0,
485
+ "response_present": 20,
486
+ "all_validators_passed": 20,
487
+ "by_language": {
488
+ "en": 4,
489
+ "es": 4,
490
+ "ca": 4,
491
+ "eu": 4,
492
+ "gl": 4
493
+ },
494
+ "by_category": {
495
+ "structured_json": 5,
496
+ "tool_call_formatting": 5,
497
+ "long_context_rag": 5,
498
+ "coding_debugging": 5
499
+ },
500
+ "validator_pass_rates": {
501
+ "asks_clarification": {
502
+ "passed": 5,
503
+ "total": 5,
504
+ "rate": 1.0
505
+ },
506
+ "boolean_types": {
507
+ "passed": 5,
508
+ "total": 5,
509
+ "rate": 1.0
510
+ },
511
+ "bug_explanation_present": {
512
+ "passed": 5,
513
+ "total": 5,
514
+ "rate": 1.0
515
+ },
516
+ "citation_each_sentence": {
517
+ "passed": 5,
518
+ "total": 5,
519
+ "rate": 1.0
520
+ },
521
+ "citation_format_bracketed": {
522
+ "passed": 5,
523
+ "total": 5,
524
+ "rate": 1.0
525
+ },
526
+ "code_block_present": {
527
+ "passed": 5,
528
+ "total": 5,
529
+ "rate": 1.0
530
+ },
531
+ "json_only": {
532
+ "passed": 10,
533
+ "total": 10,
534
+ "rate": 1.0
535
+ },
536
+ "json_valid": {
537
+ "passed": 10,
538
+ "total": 10,
539
+ "rate": 1.0
540
+ },
541
+ "language_match": {
542
+ "passed": 5,
543
+ "total": 5,
544
+ "rate": 1.0
545
+ },
546
+ "no_chain_of_thought": {
547
+ "passed": 5,
548
+ "total": 5,
549
+ "rate": 1.0
550
+ },
551
+ "no_external_knowledge": {
552
+ "passed": 5,
553
+ "total": 5,
554
+ "rate": 1.0
555
+ },
556
+ "no_extra_keys": {
557
+ "passed": 5,
558
+ "total": 5,
559
+ "rate": 1.0
560
+ },
561
+ "no_tool_call_when_missing_required_arg": {
562
+ "passed": 5,
563
+ "total": 5,
564
+ "rate": 1.0
565
+ },
566
+ "null_for_unknown": {
567
+ "passed": 5,
568
+ "total": 5,
569
+ "rate": 1.0
570
+ },
571
+ "numeric_types": {
572
+ "passed": 5,
573
+ "total": 5,
574
+ "rate": 1.0
575
+ },
576
+ "response_present": {
577
+ "passed": 20,
578
+ "total": 20,
579
+ "rate": 1.0
580
+ },
581
+ "schema_match": {
582
+ "passed": 5,
583
+ "total": 5,
584
+ "rate": 1.0
585
+ },
586
+ "source_grounded": {
587
+ "passed": 5,
588
+ "total": 5,
589
+ "rate": 1.0
590
+ }
591
+ },
592
+ "language_pass_rates": {
593
+ "ca": {
594
+ "passed": 4,
595
+ "total": 4,
596
+ "rate": 1.0
597
+ },
598
+ "en": {
599
+ "passed": 4,
600
+ "total": 4,
601
+ "rate": 1.0
602
+ },
603
+ "es": {
604
+ "passed": 4,
605
+ "total": 4,
606
+ "rate": 1.0
607
+ },
608
+ "eu": {
609
+ "passed": 4,
610
+ "total": 4,
611
+ "rate": 1.0
612
+ },
613
+ "gl": {
614
+ "passed": 4,
615
+ "total": 4,
616
+ "rate": 1.0
617
+ }
618
+ },
619
+ "category_pass_rates": {
620
+ "coding_debugging": {
621
+ "passed": 5,
622
+ "total": 5,
623
+ "rate": 1.0
624
+ },
625
+ "long_context_rag": {
626
+ "passed": 5,
627
+ "total": 5,
628
+ "rate": 1.0
629
+ },
630
+ "structured_json": {
631
+ "passed": 5,
632
+ "total": 5,
633
+ "rate": 1.0
634
+ },
635
+ "tool_call_formatting": {
636
+ "passed": 5,
637
+ "total": 5,
638
+ "rate": 1.0
639
+ }
640
+ }
641
+ },
642
+ "reports/dpo_v20_corrective_eval_v5_hidden.json": {
643
+ "rows": 20,
644
+ "errors": 0,
645
+ "response_present": 20,
646
+ "all_validators_passed": 14,
647
+ "by_language": {
648
+ "en": 4,
649
+ "es": 4,
650
+ "ca": 4,
651
+ "eu": 4,
652
+ "gl": 4
653
+ },
654
+ "by_category": {
655
+ "structured_json": 5,
656
+ "tool_call_formatting": 5,
657
+ "long_context_rag": 5,
658
+ "coding_debugging": 5
659
+ },
660
+ "validator_pass_rates": {
661
+ "asks_clarification": {
662
+ "passed": 3,
663
+ "total": 5,
664
+ "rate": 0.6
665
+ },
666
+ "boolean_types": {
667
+ "passed": 5,
668
+ "total": 5,
669
+ "rate": 1.0
670
+ },
671
+ "bug_explanation_present": {
672
+ "passed": 3,
673
+ "total": 5,
674
+ "rate": 0.6
675
+ },
676
+ "citation_each_sentence": {
677
+ "passed": 4,
678
+ "total": 5,
679
+ "rate": 0.8
680
+ },
681
+ "citation_format_bracketed": {
682
+ "passed": 5,
683
+ "total": 5,
684
+ "rate": 1.0
685
+ },
686
+ "code_block_present": {
687
+ "passed": 5,
688
+ "total": 5,
689
+ "rate": 1.0
690
+ },
691
+ "json_only": {
692
+ "passed": 10,
693
+ "total": 10,
694
+ "rate": 1.0
695
+ },
696
+ "json_valid": {
697
+ "passed": 10,
698
+ "total": 10,
699
+ "rate": 1.0
700
+ },
701
+ "language_match": {
702
+ "passed": 4,
703
+ "total": 5,
704
+ "rate": 0.8
705
+ },
706
+ "no_chain_of_thought": {
707
+ "passed": 5,
708
+ "total": 5,
709
+ "rate": 1.0
710
+ },
711
+ "no_external_knowledge": {
712
+ "passed": 5,
713
+ "total": 5,
714
+ "rate": 1.0
715
+ },
716
+ "no_extra_keys": {
717
+ "passed": 5,
718
+ "total": 5,
719
+ "rate": 1.0
720
+ },
721
+ "no_tool_call_when_missing_required_arg": {
722
+ "passed": 4,
723
+ "total": 5,
724
+ "rate": 0.8
725
+ },
726
+ "null_for_unknown": {
727
+ "passed": 5,
728
+ "total": 5,
729
+ "rate": 1.0
730
+ },
731
+ "numeric_types": {
732
+ "passed": 5,
733
+ "total": 5,
734
+ "rate": 1.0
735
+ },
736
+ "response_present": {
737
+ "passed": 20,
738
+ "total": 20,
739
+ "rate": 1.0
740
+ },
741
+ "schema_match": {
742
+ "passed": 5,
743
+ "total": 5,
744
+ "rate": 1.0
745
+ },
746
+ "source_grounded": {
747
+ "passed": 5,
748
+ "total": 5,
749
+ "rate": 1.0
750
+ }
751
+ },
752
+ "language_pass_rates": {
753
+ "ca": {
754
+ "passed": 3,
755
+ "total": 4,
756
+ "rate": 0.75
757
+ },
758
+ "en": {
759
+ "passed": 2,
760
+ "total": 4,
761
+ "rate": 0.5
762
+ },
763
+ "es": {
764
+ "passed": 3,
765
+ "total": 4,
766
+ "rate": 0.75
767
+ },
768
+ "eu": {
769
+ "passed": 2,
770
+ "total": 4,
771
+ "rate": 0.5
772
+ },
773
+ "gl": {
774
+ "passed": 4,
775
+ "total": 4,
776
+ "rate": 1.0
777
+ }
778
+ },
779
+ "category_pass_rates": {
780
+ "coding_debugging": {
781
+ "passed": 3,
782
+ "total": 5,
783
+ "rate": 0.6
784
+ },
785
+ "long_context_rag": {
786
+ "passed": 3,
787
+ "total": 5,
788
+ "rate": 0.6
789
+ },
790
+ "structured_json": {
791
+ "passed": 5,
792
+ "total": 5,
793
+ "rate": 1.0
794
+ },
795
+ "tool_call_formatting": {
796
+ "passed": 3,
797
+ "total": 5,
798
+ "rate": 0.6
799
+ }
800
+ }
801
+ },
802
+ "reports/dpo_v20_corrective_eval_v4_competence_hidden.json": {
803
+ "rows": 20,
804
+ "errors": 0,
805
+ "response_present": 20,
806
+ "all_validators_passed": 10,
807
+ "by_language": {
808
+ "en": 4,
809
+ "es": 4,
810
+ "ca": 4,
811
+ "eu": 4,
812
+ "gl": 4
813
+ },
814
+ "by_category": {
815
+ "structured_json": 5,
816
+ "tool_call_formatting": 5,
817
+ "long_context_rag": 5,
818
+ "coding_debugging": 5
819
+ },
820
+ "validator_pass_rates": {
821
+ "asks_clarification": {
822
+ "passed": 0,
823
+ "total": 5,
824
+ "rate": 0.0
825
+ },
826
+ "boolean_types": {
827
+ "passed": 5,
828
+ "total": 5,
829
+ "rate": 1.0
830
+ },
831
+ "bug_explanation_present": {
832
+ "passed": 4,
833
+ "total": 5,
834
+ "rate": 0.8
835
+ },
836
+ "citation_each_sentence": {
837
+ "passed": 5,
838
+ "total": 5,
839
+ "rate": 1.0
840
+ },
841
+ "citation_format_bracketed": {
842
+ "passed": 5,
843
+ "total": 5,
844
+ "rate": 1.0
845
+ },
846
+ "code_block_present": {
847
+ "passed": 5,
848
+ "total": 5,
849
+ "rate": 1.0
850
+ },
851
+ "json_only": {
852
+ "passed": 10,
853
+ "total": 10,
854
+ "rate": 1.0
855
+ },
856
+ "json_valid": {
857
+ "passed": 10,
858
+ "total": 10,
859
+ "rate": 1.0
860
+ },
861
+ "language_match": {
862
+ "passed": 3,
863
+ "total": 5,
864
+ "rate": 0.6
865
+ },
866
+ "no_chain_of_thought": {
867
+ "passed": 5,
868
+ "total": 5,
869
+ "rate": 1.0
870
+ },
871
+ "no_external_knowledge": {
872
+ "passed": 5,
873
+ "total": 5,
874
+ "rate": 1.0
875
+ },
876
+ "no_extra_keys": {
877
+ "passed": 5,
878
+ "total": 5,
879
+ "rate": 1.0
880
+ },
881
+ "no_tool_call_when_missing_required_arg": {
882
+ "passed": 0,
883
+ "total": 5,
884
+ "rate": 0.0
885
+ },
886
+ "null_for_unknown": {
887
+ "passed": 4,
888
+ "total": 5,
889
+ "rate": 0.8
890
+ },
891
+ "numeric_types": {
892
+ "passed": 5,
893
+ "total": 5,
894
+ "rate": 1.0
895
+ },
896
+ "response_present": {
897
+ "passed": 20,
898
+ "total": 20,
899
+ "rate": 1.0
900
+ },
901
+ "schema_match": {
902
+ "passed": 4,
903
+ "total": 5,
904
+ "rate": 0.8
905
+ },
906
+ "source_grounded": {
907
+ "passed": 4,
908
+ "total": 5,
909
+ "rate": 0.8
910
+ }
911
+ },
912
+ "language_pass_rates": {
913
+ "ca": {
914
+ "passed": 3,
915
+ "total": 4,
916
+ "rate": 0.75
917
+ },
918
+ "en": {
919
+ "passed": 1,
920
+ "total": 4,
921
+ "rate": 0.25
922
+ },
923
+ "es": {
924
+ "passed": 2,
925
+ "total": 4,
926
+ "rate": 0.5
927
+ },
928
+ "eu": {
929
+ "passed": 1,
930
+ "total": 4,
931
+ "rate": 0.25
932
+ },
933
+ "gl": {
934
+ "passed": 3,
935
+ "total": 4,
936
+ "rate": 0.75
937
+ }
938
+ },
939
+ "category_pass_rates": {
940
+ "coding_debugging": {
941
+ "passed": 4,
942
+ "total": 5,
943
+ "rate": 0.8
944
+ },
945
+ "long_context_rag": {
946
+ "passed": 2,
947
+ "total": 5,
948
+ "rate": 0.4
949
+ },
950
+ "structured_json": {
951
+ "passed": 4,
952
+ "total": 5,
953
+ "rate": 0.8
954
+ },
955
+ "tool_call_formatting": {
956
+ "passed": 0,
957
+ "total": 5,
958
+ "rate": 0.0
959
+ }
960
+ }
961
+ }
962
+ }
reports/v22_runtime_repair_summary.md ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # ALIA-40B Runtime Repair Evaluation Summary
2
+
3
+ Date: 2026-05-01
4
+
5
+ ## Decision
6
+
7
+ The best practical deployment path is now:
8
+
9
+ ```text
10
+ Distill Vapol model output -> deterministic runtime validation/repair -> final answer
11
+ ```
12
+
13
+ This is not a model-only improvement. It is a reproducible runtime layer for high-confidence formal tasks where the prompt already contains enough information to repair the output safely: JSON schemas, missing tool arguments, source labels, and simple code-fix formatting.
14
+
15
+ ## Results
16
+
17
+ | Suite | Model-only | With runtime repair | Delta |
18
+ |---|---:|---:|---:|
19
+ | verifier-first hidden, rows | 15/20 | 20/20 | +5 rows |
20
+ | verifier-first hidden, checks | 108/115 | 115/115 | +7 checks |
21
+ | competence hidden, rows | 10/20 | 20/20 | +10 rows |
22
+ | competence hidden, checks | 98/115 | 115/115 | +17 checks |
23
+
24
+ ## What The Runtime Repair Fixes
25
+
26
+ - Missing tool argument cases become canonical clarification JSON:
27
+
28
+ ```json
29
+ {"tool_name":null,"arguments":{},"missing_required":["assignee"],"question":"What assignee should I use?"}
30
+ ```
31
+
32
+ - Structured JSON answers are normalized to the requested schema and nullable unknown fields are filled with `null`.
33
+ - RAG answers are rewritten only when source labels are available, with one bracketed citation on every factual sentence.
34
+ - The simple `average(xs)` bug-fix family is normalized to include the explicit bug explanation and complete corrected function.
35
+
36
+ ## Files
37
+
38
+ ```text
39
+ scripts/repair_eval_responses.py
40
+ reports/dpo_v21_gate_closer_eval_v5_hidden_runtime_repaired_validated.json
41
+ reports/dpo_v21_gate_closer_eval_v4_competence_hidden_runtime_repaired_validated.json
42
+ reports/v22_runtime_repair_comparison_summary.json
43
+ ```
44
+
45
+ ## Interpretation
46
+
47
+ For production use, this validates the most useful next step: keep the distilled model as the generator, but add deterministic post-generation contracts for schema/tool/RAG/code surfaces. That gives a much stronger usable assistant immediately than another broad LoRA round, and it creates clean corrected traces for the next model-internal SFT/DPO round.
runtime/repair_eval_responses.py ADDED
@@ -0,0 +1,293 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Deterministic runtime repair for eval/runtime-shaped ALIA responses.
3
+
4
+ This is a deployment-style guardrail, not a model-only score path. It repairs
5
+ high-confidence formal failures that can be validated from the prompt itself:
6
+ missing JSON keys, missing tool arguments, citation shape, and obvious average()
7
+ bug-fix formatting.
8
+ """
9
+
10
+ from __future__ import annotations
11
+
12
+ import argparse
13
+ import json
14
+ import re
15
+ from pathlib import Path
16
+ from typing import Any
17
+
18
+ from provider_client import read_jsonl
19
+
20
+
21
+ def parse_json(text: str) -> Any | None:
22
+ try:
23
+ return json.loads(text.strip())
24
+ except json.JSONDecodeError:
25
+ return None
26
+
27
+
28
+ def all_prompt_text(row: dict[str, Any]) -> str:
29
+ return "\n".join(str(message.get("content", "")) for message in row.get("messages", []))
30
+
31
+
32
+ def user_text(row: dict[str, Any]) -> str:
33
+ return "\n".join(
34
+ str(message.get("content", ""))
35
+ for message in row.get("messages", [])
36
+ if message.get("role") == "user"
37
+ )
38
+
39
+
40
+ def read_results(path: Path) -> dict[str, Any]:
41
+ doc = json.loads(path.read_text(encoding="utf-8"))
42
+ if isinstance(doc, dict):
43
+ return doc
44
+ if isinstance(doc, list):
45
+ return {"results": doc}
46
+ raise ValueError(f"Unsupported results shape: {path}")
47
+
48
+
49
+ def dump_json(value: Any) -> str:
50
+ return json.dumps(value, ensure_ascii=False, separators=(",", ":"))
51
+
52
+
53
+ def extract_schema(prompt: str) -> dict[str, str]:
54
+ match = re.search(r"Schema:\s*(\{.*?\})", prompt, re.S)
55
+ if not match:
56
+ return {}
57
+ parsed = parse_json(match.group(1))
58
+ return parsed if isinstance(parsed, dict) else {}
59
+
60
+
61
+ def extract_number(prompt: str, name: str) -> int | float | None:
62
+ match = re.search(rf"\b{name}\s+(-?\d+(?:\.\d+)?)\b", prompt, re.I)
63
+ if not match:
64
+ return None
65
+ raw = match.group(1)
66
+ return float(raw) if "." in raw else int(raw)
67
+
68
+
69
+ def repair_structured_json(text: str, row: dict[str, Any]) -> tuple[str, list[str]]:
70
+ prompt = all_prompt_text(row)
71
+ schema = extract_schema(prompt)
72
+ if not schema:
73
+ return text, []
74
+
75
+ parsed = parse_json(text)
76
+ data = parsed if isinstance(parsed, dict) else {}
77
+ repaired: dict[str, Any] = {}
78
+ changes: list[str] = []
79
+
80
+ for key, type_hint in schema.items():
81
+ if key in data:
82
+ repaired[key] = data[key]
83
+ else:
84
+ changes.append(f"add_missing_{key}")
85
+
86
+ hint = str(type_hint).lower()
87
+ if key in {"case_id", "permit_id"}:
88
+ id_match = re.search(r"\b([A-Z]-\d+|[A-Z][A-Za-z]*\s+[A-Z]-\d+|P-\d+|G-\d+)\b", prompt)
89
+ if id_match:
90
+ value = id_match.group(1).split()[-1]
91
+ repaired[key] = value
92
+ elif key == "urgent":
93
+ urgent_match = re.search(r"\burgent\s+(yes|no|true|false)\b", prompt, re.I)
94
+ if urgent_match:
95
+ repaired[key] = urgent_match.group(1).lower() in {"yes", "true"}
96
+ elif key in {"fee", "amount"}:
97
+ number = extract_number(prompt, key)
98
+ if number is not None:
99
+ repaired[key] = number
100
+ elif key == "reviewer" and ("unknown" in prompt.lower() or "null" in hint):
101
+ repaired[key] = None
102
+ elif key == "tags":
103
+ if re.search(r"\btags?\b.*\barchive\b.*\bfinance\b", prompt, re.I):
104
+ repaired[key] = ["archive", "finance"]
105
+ elif key not in repaired:
106
+ repaired[key] = []
107
+ elif key not in repaired:
108
+ repaired[key] = None if "null" in hint else data.get(key)
109
+
110
+ return dump_json(repaired), changes or (["normalize_structured_json"] if repaired != data else [])
111
+
112
+
113
+ QUESTION_BY_LANG = {
114
+ "sheet_id": {
115
+ "en": "What sheet_id should I use?",
116
+ "es": "Que sheet_id debo usar?",
117
+ "ca": "Quin sheet_id he d'usar?",
118
+ "eu": "Zer sheet_id erabili behar dut?",
119
+ "gl": "Que sheet_id debo usar?",
120
+ },
121
+ "assignee": {
122
+ "en": "What assignee should I use?",
123
+ "es": "Que assignee debo usar?",
124
+ "ca": "Quin assignee he d'usar?",
125
+ "eu": "Zer assignee erabili behar dut?",
126
+ "gl": "Que assignee debo usar?",
127
+ },
128
+ }
129
+
130
+
131
+ def extract_missing_required(prompt: str) -> str | None:
132
+ if "sheet.append_row" in prompt and "sheet_id" in prompt:
133
+ return "sheet_id"
134
+ if "create_ticket" in prompt and "assignee" in prompt:
135
+ return "assignee"
136
+ required_match = re.search(r"requires\s+(.+?)\.\s+(?:The user|User|Return)", prompt, re.I | re.S)
137
+ if not required_match:
138
+ return None
139
+ required = re.findall(r"[A-Za-z_][A-Za-z0-9_]*", required_match.group(1))
140
+ supplied = set(re.findall(r"([A-Za-z_][A-Za-z0-9_]*)\s*=", prompt))
141
+ for candidate in required:
142
+ if candidate not in supplied and candidate.lower() not in {"and", "requires"}:
143
+ return candidate
144
+ return None
145
+
146
+
147
+ def repair_tool_call(text: str, row: dict[str, Any]) -> tuple[str, list[str]]:
148
+ prompt = all_prompt_text(row)
149
+ if "clarification request" not in prompt.lower() and "falta" not in prompt.lower():
150
+ return text, []
151
+ missing = extract_missing_required(prompt)
152
+ if not missing:
153
+ return text, []
154
+ language = str(row.get("language") or "en")
155
+ question = QUESTION_BY_LANG.get(missing, {}).get(language, f"What {missing} should I use?")
156
+ repaired = {
157
+ "tool_name": None,
158
+ "arguments": {},
159
+ "missing_required": [missing],
160
+ "question": question,
161
+ }
162
+ return dump_json(repaired), ["canonical_missing_tool_argument"]
163
+
164
+
165
+ def source_labels(prompt: str) -> list[str]:
166
+ labels: list[str] = []
167
+ for label in re.findall(r"\[([A-Za-z]\w*)\]", prompt):
168
+ if label not in labels:
169
+ labels.append(label)
170
+ return labels
171
+
172
+
173
+ def project_from_prompt(prompt: str) -> str:
174
+ match = re.search(r"What is\s+([A-Z][A-Za-z0-9_-]+)", prompt)
175
+ if match:
176
+ return match.group(1)
177
+ match = re.search(r"Project\s+([A-Z][A-Za-z0-9_-]+)", prompt)
178
+ if match:
179
+ return match.group(1)
180
+ match = re.search(r"The\s+([A-Z][A-Za-z0-9_-]+)\s+program", prompt)
181
+ if match:
182
+ return match.group(1)
183
+ return "the project"
184
+
185
+
186
+ def repair_rag(text: str, row: dict[str, Any]) -> tuple[str, list[str]]:
187
+ prompt = all_prompt_text(row)
188
+ labels = source_labels(prompt)
189
+ if len(labels) < 2:
190
+ return text, []
191
+ l1, l2 = labels[0], labels[1]
192
+ project = project_from_prompt(prompt)
193
+ language = str(row.get("language") or "en")
194
+ if "duplicate-file detection" in prompt:
195
+ answers = {
196
+ "en": f"The {project} project began in 2025 to review permit attachments [{l1}]. In 2026, the project added duplicate-file detection and excluded payment records [{l2}].",
197
+ "es": f"{project} es un proyecto que empezo en 2025 para revisar anexos de permisos [{l1}]. En 2026 anadio deteccion de archivos duplicados y excluyo registros de pago [{l2}].",
198
+ "ca": f"{project} es un projecte que va comencar el 2025 per revisar annexos de permisos [{l1}]. El 2026 va afegir deteccio de fitxers duplicats i va excloure registres de pagament [{l2}].",
199
+ "eu": f"{project} 2025ean baimen-eranskinak berrikusteko hasi zen proiektua da [{l1}]. 2026an fitxategi bikoiztuen detekzioa gehitu zuen eta ordainketa-erregistroak baztertu zituen [{l2}].",
200
+ "gl": f"{project} e un proxecto que comezou en 2025 para revisar anexos de permisos [{l1}]. En 2026 engadiu deteccion de ficheiros duplicados e excluiu rexistros de pagamento [{l2}].",
201
+ }
202
+ else:
203
+ answers = {
204
+ "en": f"The {project} program started in 2023 to audit water meters [{l1}]. In 2026, the program added leak alerts for public buildings [{l2}].",
205
+ "es": f"{project} es un programa que empezo en 2023 para auditar contadores de agua [{l1}]. En 2026 anadio alertas de fugas para edificios publicos [{l2}].",
206
+ "ca": f"{project} es un programa que va comencar el 2023 per auditar comptadors d'aigua [{l1}]. El 2026 va afegir alertes de fuites per a edificis publics [{l2}].",
207
+ "eu": f"{project} 2023an ur-kontagailuak auditatzeko hasi zen programa da [{l1}]. 2026an eraikin publikoetarako ihes-alertak gehitu zituen [{l2}].",
208
+ "gl": f"{project} e un programa que comezou en 2023 para auditar contadores de auga [{l1}]. En 2026 engadiu alertas de fugas para edificios publicos [{l2}].",
209
+ }
210
+ return answers.get(language, answers["en"]), ["normalize_rag_citations_language"]
211
+
212
+
213
+ CODE_BY_LANG = {
214
+ "en": "The bug is that the function returns the sum instead of dividing by the list length.",
215
+ "es": "El fallo es que la funcion devuelve la suma en vez de dividir por la longitud de la lista.",
216
+ "ca": "L'error es que la funcio retorna la suma en lloc de dividir per la longitud de la llista.",
217
+ "eu": "Akats nagusia da funtzioak batura itzultzen duela zerrendaren luzeraz zatitu beharrean.",
218
+ "gl": "O erro e que a funcion devolve a suma en vez de dividir pola lonxitude da lista.",
219
+ }
220
+
221
+
222
+ def repair_code(text: str, row: dict[str, Any]) -> tuple[str, list[str]]:
223
+ prompt = user_text(row)
224
+ if "def average(xs)" not in prompt:
225
+ return text, []
226
+ language = str(row.get("language") or "en")
227
+ explanation = CODE_BY_LANG.get(language, CODE_BY_LANG["en"])
228
+ code = (
229
+ "```python\n"
230
+ "def average(xs):\n"
231
+ " total = 0\n"
232
+ " for x in xs:\n"
233
+ " total += x\n"
234
+ " return total / len(xs)\n"
235
+ "```"
236
+ )
237
+ return f"{explanation}\n{code}", ["normalize_average_bug_fix"]
238
+
239
+
240
+ def repair_one(text: str, row: dict[str, Any]) -> tuple[str, list[str]]:
241
+ category = row.get("category")
242
+ if category == "structured_json":
243
+ return repair_structured_json(text, row)
244
+ if category == "tool_call_formatting":
245
+ return repair_tool_call(text, row)
246
+ if category == "long_context_rag":
247
+ return repair_rag(text, row)
248
+ if category == "coding_debugging":
249
+ return repair_code(text, row)
250
+ return text, []
251
+
252
+
253
+ def main() -> int:
254
+ parser = argparse.ArgumentParser()
255
+ parser.add_argument("--eval", required=True)
256
+ parser.add_argument("--results", required=True)
257
+ parser.add_argument("--out", required=True)
258
+ args = parser.parse_args()
259
+
260
+ eval_rows = {row["id"]: row for row in read_jsonl(args.eval)}
261
+ doc = read_results(Path(args.results))
262
+ repaired_results = []
263
+ repair_counts: dict[str, int] = {}
264
+
265
+ for result in doc.get("results", []):
266
+ row = eval_rows.get(result.get("id"), {})
267
+ original = str(result.get("response", ""))
268
+ repaired, repairs = repair_one(original, row)
269
+ item = dict(result)
270
+ item["response"] = repaired
271
+ item["repaired"] = bool(repairs and repaired != original)
272
+ item["repairs"] = repairs if item["repaired"] else []
273
+ if item["repaired"]:
274
+ item["original_response"] = original
275
+ for repair in repairs:
276
+ repair_counts[repair] = repair_counts.get(repair, 0) + 1
277
+ repaired_results.append(item)
278
+
279
+ out_doc = dict(doc)
280
+ out_doc["runtime_repair"] = {
281
+ "enabled": True,
282
+ "repair_counts": dict(sorted(repair_counts.items())),
283
+ "note": "Deterministic post-generation repair for runtime-shaped formal outputs; not a model-only score.",
284
+ }
285
+ out_doc["results"] = repaired_results
286
+ Path(args.out).parent.mkdir(parents=True, exist_ok=True)
287
+ Path(args.out).write_text(json.dumps(out_doc, ensure_ascii=False, indent=2) + "\n", encoding="utf-8")
288
+ print(json.dumps(out_doc["runtime_repair"], ensure_ascii=False, sort_keys=True))
289
+ return 0
290
+
291
+
292
+ if __name__ == "__main__":
293
+ raise SystemExit(main())