thuan220401 commited on
Commit
06be6b4
·
verified ·
1 Parent(s): f401052

Delete checkpoint-250

Browse files
checkpoint-250/README.md DELETED
@@ -1,202 +0,0 @@
1
- ---
2
- base_model: NousResearch/Hermes-3-Llama-3.1-8B
3
- library_name: peft
4
- ---
5
-
6
- # Model Card for Model ID
7
-
8
- <!-- Provide a quick summary of what the model is/does. -->
9
-
10
-
11
-
12
- ## Model Details
13
-
14
- ### Model Description
15
-
16
- <!-- Provide a longer summary of what this model is. -->
17
-
18
-
19
-
20
- - **Developed by:** [More Information Needed]
21
- - **Funded by [optional]:** [More Information Needed]
22
- - **Shared by [optional]:** [More Information Needed]
23
- - **Model type:** [More Information Needed]
24
- - **Language(s) (NLP):** [More Information Needed]
25
- - **License:** [More Information Needed]
26
- - **Finetuned from model [optional]:** [More Information Needed]
27
-
28
- ### Model Sources [optional]
29
-
30
- <!-- Provide the basic links for the model. -->
31
-
32
- - **Repository:** [More Information Needed]
33
- - **Paper [optional]:** [More Information Needed]
34
- - **Demo [optional]:** [More Information Needed]
35
-
36
- ## Uses
37
-
38
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
39
-
40
- ### Direct Use
41
-
42
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
43
-
44
- [More Information Needed]
45
-
46
- ### Downstream Use [optional]
47
-
48
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
49
-
50
- [More Information Needed]
51
-
52
- ### Out-of-Scope Use
53
-
54
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
55
-
56
- [More Information Needed]
57
-
58
- ## Bias, Risks, and Limitations
59
-
60
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
61
-
62
- [More Information Needed]
63
-
64
- ### Recommendations
65
-
66
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
67
-
68
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
69
-
70
- ## How to Get Started with the Model
71
-
72
- Use the code below to get started with the model.
73
-
74
- [More Information Needed]
75
-
76
- ## Training Details
77
-
78
- ### Training Data
79
-
80
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
81
-
82
- [More Information Needed]
83
-
84
- ### Training Procedure
85
-
86
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
87
-
88
- #### Preprocessing [optional]
89
-
90
- [More Information Needed]
91
-
92
-
93
- #### Training Hyperparameters
94
-
95
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
96
-
97
- #### Speeds, Sizes, Times [optional]
98
-
99
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
100
-
101
- [More Information Needed]
102
-
103
- ## Evaluation
104
-
105
- <!-- This section describes the evaluation protocols and provides the results. -->
106
-
107
- ### Testing Data, Factors & Metrics
108
-
109
- #### Testing Data
110
-
111
- <!-- This should link to a Dataset Card if possible. -->
112
-
113
- [More Information Needed]
114
-
115
- #### Factors
116
-
117
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
118
-
119
- [More Information Needed]
120
-
121
- #### Metrics
122
-
123
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
124
-
125
- [More Information Needed]
126
-
127
- ### Results
128
-
129
- [More Information Needed]
130
-
131
- #### Summary
132
-
133
-
134
-
135
- ## Model Examination [optional]
136
-
137
- <!-- Relevant interpretability work for the model goes here -->
138
-
139
- [More Information Needed]
140
-
141
- ## Environmental Impact
142
-
143
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
144
-
145
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
146
-
147
- - **Hardware Type:** [More Information Needed]
148
- - **Hours used:** [More Information Needed]
149
- - **Cloud Provider:** [More Information Needed]
150
- - **Compute Region:** [More Information Needed]
151
- - **Carbon Emitted:** [More Information Needed]
152
-
153
- ## Technical Specifications [optional]
154
-
155
- ### Model Architecture and Objective
156
-
157
- [More Information Needed]
158
-
159
- ### Compute Infrastructure
160
-
161
- [More Information Needed]
162
-
163
- #### Hardware
164
-
165
- [More Information Needed]
166
-
167
- #### Software
168
-
169
- [More Information Needed]
170
-
171
- ## Citation [optional]
172
-
173
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
174
-
175
- **BibTeX:**
176
-
177
- [More Information Needed]
178
-
179
- **APA:**
180
-
181
- [More Information Needed]
182
-
183
- ## Glossary [optional]
184
-
185
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
186
-
187
- [More Information Needed]
188
-
189
- ## More Information [optional]
190
-
191
- [More Information Needed]
192
-
193
- ## Model Card Authors [optional]
194
-
195
- [More Information Needed]
196
-
197
- ## Model Card Contact
198
-
199
- [More Information Needed]
200
- ### Framework versions
201
-
202
- - PEFT 0.13.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
checkpoint-250/adapter_config.json DELETED
@@ -1,34 +0,0 @@
1
- {
2
- "alpha_pattern": {},
3
- "auto_mapping": null,
4
- "base_model_name_or_path": "NousResearch/Hermes-3-Llama-3.1-8B",
5
- "bias": "none",
6
- "fan_in_fan_out": false,
7
- "inference_mode": true,
8
- "init_lora_weights": true,
9
- "layer_replication": null,
10
- "layers_pattern": null,
11
- "layers_to_transform": null,
12
- "loftq_config": {},
13
- "lora_alpha": 64,
14
- "lora_dropout": 0.05,
15
- "megatron_config": null,
16
- "megatron_core": "megatron.core",
17
- "modules_to_save": null,
18
- "peft_type": "LORA",
19
- "r": 16,
20
- "rank_pattern": {},
21
- "revision": null,
22
- "target_modules": [
23
- "k_proj",
24
- "v_proj",
25
- "up_proj",
26
- "down_proj",
27
- "q_proj",
28
- "o_proj",
29
- "gate_proj"
30
- ],
31
- "task_type": "CAUSAL_LM",
32
- "use_dora": true,
33
- "use_rslora": false
34
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
checkpoint-250/adapter_model.safetensors DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:8451ebbf7c7c3cc19995b8edbf2e9b0c8a8a8aa30dd3af48723b54628f960751
3
- size 173368456
 
 
 
 
checkpoint-250/optimizer.pt DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:8e10079989dcc361ae722f07fa8043bc1448b23156d48cfab059e0f9fd16bc48
3
- size 347119730
 
 
 
 
checkpoint-250/rng_state.pth DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:e5e616e9dbd425e18bd7c1d6d00e918c5a6ff17a14957e778697e7854b7a597c
3
- size 14244
 
 
 
 
checkpoint-250/scheduler.pt DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:1d1650f5062195d8ee65b24ab00a137ab48cccbff41f41ba060d4208547a763c
3
- size 1064
 
 
 
 
checkpoint-250/special_tokens_map.json DELETED
@@ -1,21 +0,0 @@
1
- {
2
- "additional_special_tokens": [
3
- {
4
- "content": "<|im_start|>",
5
- "lstrip": false,
6
- "normalized": false,
7
- "rstrip": false,
8
- "single_word": false
9
- },
10
- {
11
- "content": "<|im_end|>",
12
- "lstrip": false,
13
- "normalized": false,
14
- "rstrip": false,
15
- "single_word": false
16
- }
17
- ],
18
- "bos_token": "<|im_start|>",
19
- "eos_token": "<|im_end|>",
20
- "pad_token": "<|im_end|>"
21
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
checkpoint-250/tokenizer.json DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:038968f16122fee53c1cd8b410a01c21d9cf995cda80be434bb34e4376820027
3
- size 17209500
 
 
 
 
checkpoint-250/tokenizer_config.json DELETED
@@ -1,2067 +0,0 @@
1
- {
2
- "added_tokens_decoder": {
3
- "128000": {
4
- "content": "<|begin_of_text|>",
5
- "lstrip": false,
6
- "normalized": false,
7
- "rstrip": false,
8
- "single_word": false,
9
- "special": true
10
- },
11
- "128001": {
12
- "content": "<|end_of_text|>",
13
- "lstrip": false,
14
- "normalized": false,
15
- "rstrip": false,
16
- "single_word": false,
17
- "special": true
18
- },
19
- "128002": {
20
- "content": "<tool_call>",
21
- "lstrip": false,
22
- "normalized": false,
23
- "rstrip": false,
24
- "single_word": false,
25
- "special": false
26
- },
27
- "128003": {
28
- "content": "<tool_response>",
29
- "lstrip": false,
30
- "normalized": false,
31
- "rstrip": false,
32
- "single_word": false,
33
- "special": false
34
- },
35
- "128004": {
36
- "content": "<|finetune_right_pad_id|>",
37
- "lstrip": false,
38
- "normalized": false,
39
- "rstrip": false,
40
- "single_word": false,
41
- "special": true
42
- },
43
- "128005": {
44
- "content": "<|reserved_special_token_2|>",
45
- "lstrip": false,
46
- "normalized": false,
47
- "rstrip": false,
48
- "single_word": false,
49
- "special": true
50
- },
51
- "128006": {
52
- "content": "<|start_header_id|>",
53
- "lstrip": false,
54
- "normalized": false,
55
- "rstrip": false,
56
- "single_word": false,
57
- "special": true
58
- },
59
- "128007": {
60
- "content": "<|end_header_id|>",
61
- "lstrip": false,
62
- "normalized": false,
63
- "rstrip": false,
64
- "single_word": false,
65
- "special": true
66
- },
67
- "128008": {
68
- "content": "<|eom_id|>",
69
- "lstrip": false,
70
- "normalized": false,
71
- "rstrip": false,
72
- "single_word": false,
73
- "special": true
74
- },
75
- "128009": {
76
- "content": "<|eot_id|>",
77
- "lstrip": false,
78
- "normalized": false,
79
- "rstrip": false,
80
- "single_word": false,
81
- "special": true
82
- },
83
- "128010": {
84
- "content": "<|python_tag|>",
85
- "lstrip": false,
86
- "normalized": false,
87
- "rstrip": false,
88
- "single_word": false,
89
- "special": true
90
- },
91
- "128011": {
92
- "content": "<tools>",
93
- "lstrip": false,
94
- "normalized": false,
95
- "rstrip": false,
96
- "single_word": false,
97
- "special": false
98
- },
99
- "128012": {
100
- "content": "</tools>",
101
- "lstrip": false,
102
- "normalized": false,
103
- "rstrip": false,
104
- "single_word": false,
105
- "special": false
106
- },
107
- "128013": {
108
- "content": "</tool_call>",
109
- "lstrip": false,
110
- "normalized": false,
111
- "rstrip": false,
112
- "single_word": false,
113
- "special": false
114
- },
115
- "128014": {
116
- "content": "</tool_response>",
117
- "lstrip": false,
118
- "normalized": false,
119
- "rstrip": false,
120
- "single_word": false,
121
- "special": false
122
- },
123
- "128015": {
124
- "content": "<schema>",
125
- "lstrip": false,
126
- "normalized": false,
127
- "rstrip": false,
128
- "single_word": false,
129
- "special": false
130
- },
131
- "128016": {
132
- "content": "</schema>",
133
- "lstrip": false,
134
- "normalized": false,
135
- "rstrip": false,
136
- "single_word": false,
137
- "special": false
138
- },
139
- "128017": {
140
- "content": "<scratch_pad>",
141
- "lstrip": false,
142
- "normalized": false,
143
- "rstrip": false,
144
- "single_word": false,
145
- "special": false
146
- },
147
- "128018": {
148
- "content": "</scratch_pad>",
149
- "lstrip": false,
150
- "normalized": false,
151
- "rstrip": false,
152
- "single_word": false,
153
- "special": false
154
- },
155
- "128019": {
156
- "content": "<SCRATCHPAD>",
157
- "lstrip": false,
158
- "normalized": false,
159
- "rstrip": false,
160
- "single_word": false,
161
- "special": false
162
- },
163
- "128020": {
164
- "content": "</SCRATCHPAD>",
165
- "lstrip": false,
166
- "normalized": false,
167
- "rstrip": false,
168
- "single_word": false,
169
- "special": false
170
- },
171
- "128021": {
172
- "content": "<REASONING>",
173
- "lstrip": false,
174
- "normalized": false,
175
- "rstrip": false,
176
- "single_word": false,
177
- "special": false
178
- },
179
- "128022": {
180
- "content": "</REASONING>",
181
- "lstrip": false,
182
- "normalized": false,
183
- "rstrip": false,
184
- "single_word": false,
185
- "special": false
186
- },
187
- "128023": {
188
- "content": "<INNER_MONOLOGUE>",
189
- "lstrip": false,
190
- "normalized": false,
191
- "rstrip": false,
192
- "single_word": false,
193
- "special": false
194
- },
195
- "128024": {
196
- "content": "</INNER_MONOLOGUE>",
197
- "lstrip": false,
198
- "normalized": false,
199
- "rstrip": false,
200
- "single_word": false,
201
- "special": false
202
- },
203
- "128025": {
204
- "content": "<PLAN>",
205
- "lstrip": false,
206
- "normalized": false,
207
- "rstrip": false,
208
- "single_word": false,
209
- "special": false
210
- },
211
- "128026": {
212
- "content": "</PLAN>",
213
- "lstrip": false,
214
- "normalized": false,
215
- "rstrip": false,
216
- "single_word": false,
217
- "special": false
218
- },
219
- "128027": {
220
- "content": "<EXECUTION>",
221
- "lstrip": false,
222
- "normalized": false,
223
- "rstrip": false,
224
- "single_word": false,
225
- "special": false
226
- },
227
- "128028": {
228
- "content": "</EXECUTION>",
229
- "lstrip": false,
230
- "normalized": false,
231
- "rstrip": false,
232
- "single_word": false,
233
- "special": false
234
- },
235
- "128029": {
236
- "content": "<REFLECTION>",
237
- "lstrip": false,
238
- "normalized": false,
239
- "rstrip": false,
240
- "single_word": false,
241
- "special": false
242
- },
243
- "128030": {
244
- "content": "</REFLECTION>",
245
- "lstrip": false,
246
- "normalized": false,
247
- "rstrip": false,
248
- "single_word": false,
249
- "special": false
250
- },
251
- "128031": {
252
- "content": "<THINKING>",
253
- "lstrip": false,
254
- "normalized": false,
255
- "rstrip": false,
256
- "single_word": false,
257
- "special": false
258
- },
259
- "128032": {
260
- "content": "</THINKING>",
261
- "lstrip": false,
262
- "normalized": false,
263
- "rstrip": false,
264
- "single_word": false,
265
- "special": false
266
- },
267
- "128033": {
268
- "content": "<SOLUTION>",
269
- "lstrip": false,
270
- "normalized": false,
271
- "rstrip": false,
272
- "single_word": false,
273
- "special": false
274
- },
275
- "128034": {
276
- "content": "</SOLUTION>",
277
- "lstrip": false,
278
- "normalized": false,
279
- "rstrip": false,
280
- "single_word": false,
281
- "special": false
282
- },
283
- "128035": {
284
- "content": "<EXPLANATION>",
285
- "lstrip": false,
286
- "normalized": false,
287
- "rstrip": false,
288
- "single_word": false,
289
- "special": false
290
- },
291
- "128036": {
292
- "content": "</EXPLANATION>",
293
- "lstrip": false,
294
- "normalized": false,
295
- "rstrip": false,
296
- "single_word": false,
297
- "special": false
298
- },
299
- "128037": {
300
- "content": "<UNIT_TEST>",
301
- "lstrip": false,
302
- "normalized": false,
303
- "rstrip": false,
304
- "single_word": false,
305
- "special": false
306
- },
307
- "128038": {
308
- "content": "</UNIT_TEST>",
309
- "lstrip": false,
310
- "normalized": false,
311
- "rstrip": false,
312
- "single_word": false,
313
- "special": false
314
- },
315
- "128039": {
316
- "content": "<|im_start|>",
317
- "lstrip": false,
318
- "normalized": false,
319
- "rstrip": false,
320
- "single_word": false,
321
- "special": true
322
- },
323
- "128040": {
324
- "content": "<|im_end|>",
325
- "lstrip": false,
326
- "normalized": false,
327
- "rstrip": false,
328
- "single_word": false,
329
- "special": true
330
- },
331
- "128041": {
332
- "content": "<|reserved_special_token_33|>",
333
- "lstrip": false,
334
- "normalized": false,
335
- "rstrip": false,
336
- "single_word": false,
337
- "special": true
338
- },
339
- "128042": {
340
- "content": "<|reserved_special_token_34|>",
341
- "lstrip": false,
342
- "normalized": false,
343
- "rstrip": false,
344
- "single_word": false,
345
- "special": true
346
- },
347
- "128043": {
348
- "content": "<|reserved_special_token_35|>",
349
- "lstrip": false,
350
- "normalized": false,
351
- "rstrip": false,
352
- "single_word": false,
353
- "special": true
354
- },
355
- "128044": {
356
- "content": "<|reserved_special_token_36|>",
357
- "lstrip": false,
358
- "normalized": false,
359
- "rstrip": false,
360
- "single_word": false,
361
- "special": true
362
- },
363
- "128045": {
364
- "content": "<|reserved_special_token_37|>",
365
- "lstrip": false,
366
- "normalized": false,
367
- "rstrip": false,
368
- "single_word": false,
369
- "special": true
370
- },
371
- "128046": {
372
- "content": "<|reserved_special_token_38|>",
373
- "lstrip": false,
374
- "normalized": false,
375
- "rstrip": false,
376
- "single_word": false,
377
- "special": true
378
- },
379
- "128047": {
380
- "content": "<|reserved_special_token_39|>",
381
- "lstrip": false,
382
- "normalized": false,
383
- "rstrip": false,
384
- "single_word": false,
385
- "special": true
386
- },
387
- "128048": {
388
- "content": "<|reserved_special_token_40|>",
389
- "lstrip": false,
390
- "normalized": false,
391
- "rstrip": false,
392
- "single_word": false,
393
- "special": true
394
- },
395
- "128049": {
396
- "content": "<|reserved_special_token_41|>",
397
- "lstrip": false,
398
- "normalized": false,
399
- "rstrip": false,
400
- "single_word": false,
401
- "special": true
402
- },
403
- "128050": {
404
- "content": "<|reserved_special_token_42|>",
405
- "lstrip": false,
406
- "normalized": false,
407
- "rstrip": false,
408
- "single_word": false,
409
- "special": true
410
- },
411
- "128051": {
412
- "content": "<|reserved_special_token_43|>",
413
- "lstrip": false,
414
- "normalized": false,
415
- "rstrip": false,
416
- "single_word": false,
417
- "special": true
418
- },
419
- "128052": {
420
- "content": "<|reserved_special_token_44|>",
421
- "lstrip": false,
422
- "normalized": false,
423
- "rstrip": false,
424
- "single_word": false,
425
- "special": true
426
- },
427
- "128053": {
428
- "content": "<|reserved_special_token_45|>",
429
- "lstrip": false,
430
- "normalized": false,
431
- "rstrip": false,
432
- "single_word": false,
433
- "special": true
434
- },
435
- "128054": {
436
- "content": "<|reserved_special_token_46|>",
437
- "lstrip": false,
438
- "normalized": false,
439
- "rstrip": false,
440
- "single_word": false,
441
- "special": true
442
- },
443
- "128055": {
444
- "content": "<|reserved_special_token_47|>",
445
- "lstrip": false,
446
- "normalized": false,
447
- "rstrip": false,
448
- "single_word": false,
449
- "special": true
450
- },
451
- "128056": {
452
- "content": "<|reserved_special_token_48|>",
453
- "lstrip": false,
454
- "normalized": false,
455
- "rstrip": false,
456
- "single_word": false,
457
- "special": true
458
- },
459
- "128057": {
460
- "content": "<|reserved_special_token_49|>",
461
- "lstrip": false,
462
- "normalized": false,
463
- "rstrip": false,
464
- "single_word": false,
465
- "special": true
466
- },
467
- "128058": {
468
- "content": "<|reserved_special_token_50|>",
469
- "lstrip": false,
470
- "normalized": false,
471
- "rstrip": false,
472
- "single_word": false,
473
- "special": true
474
- },
475
- "128059": {
476
- "content": "<|reserved_special_token_51|>",
477
- "lstrip": false,
478
- "normalized": false,
479
- "rstrip": false,
480
- "single_word": false,
481
- "special": true
482
- },
483
- "128060": {
484
- "content": "<|reserved_special_token_52|>",
485
- "lstrip": false,
486
- "normalized": false,
487
- "rstrip": false,
488
- "single_word": false,
489
- "special": true
490
- },
491
- "128061": {
492
- "content": "<|reserved_special_token_53|>",
493
- "lstrip": false,
494
- "normalized": false,
495
- "rstrip": false,
496
- "single_word": false,
497
- "special": true
498
- },
499
- "128062": {
500
- "content": "<|reserved_special_token_54|>",
501
- "lstrip": false,
502
- "normalized": false,
503
- "rstrip": false,
504
- "single_word": false,
505
- "special": true
506
- },
507
- "128063": {
508
- "content": "<|reserved_special_token_55|>",
509
- "lstrip": false,
510
- "normalized": false,
511
- "rstrip": false,
512
- "single_word": false,
513
- "special": true
514
- },
515
- "128064": {
516
- "content": "<|reserved_special_token_56|>",
517
- "lstrip": false,
518
- "normalized": false,
519
- "rstrip": false,
520
- "single_word": false,
521
- "special": true
522
- },
523
- "128065": {
524
- "content": "<|reserved_special_token_57|>",
525
- "lstrip": false,
526
- "normalized": false,
527
- "rstrip": false,
528
- "single_word": false,
529
- "special": true
530
- },
531
- "128066": {
532
- "content": "<|reserved_special_token_58|>",
533
- "lstrip": false,
534
- "normalized": false,
535
- "rstrip": false,
536
- "single_word": false,
537
- "special": true
538
- },
539
- "128067": {
540
- "content": "<|reserved_special_token_59|>",
541
- "lstrip": false,
542
- "normalized": false,
543
- "rstrip": false,
544
- "single_word": false,
545
- "special": true
546
- },
547
- "128068": {
548
- "content": "<|reserved_special_token_60|>",
549
- "lstrip": false,
550
- "normalized": false,
551
- "rstrip": false,
552
- "single_word": false,
553
- "special": true
554
- },
555
- "128069": {
556
- "content": "<|reserved_special_token_61|>",
557
- "lstrip": false,
558
- "normalized": false,
559
- "rstrip": false,
560
- "single_word": false,
561
- "special": true
562
- },
563
- "128070": {
564
- "content": "<|reserved_special_token_62|>",
565
- "lstrip": false,
566
- "normalized": false,
567
- "rstrip": false,
568
- "single_word": false,
569
- "special": true
570
- },
571
- "128071": {
572
- "content": "<|reserved_special_token_63|>",
573
- "lstrip": false,
574
- "normalized": false,
575
- "rstrip": false,
576
- "single_word": false,
577
- "special": true
578
- },
579
- "128072": {
580
- "content": "<|reserved_special_token_64|>",
581
- "lstrip": false,
582
- "normalized": false,
583
- "rstrip": false,
584
- "single_word": false,
585
- "special": true
586
- },
587
- "128073": {
588
- "content": "<|reserved_special_token_65|>",
589
- "lstrip": false,
590
- "normalized": false,
591
- "rstrip": false,
592
- "single_word": false,
593
- "special": true
594
- },
595
- "128074": {
596
- "content": "<|reserved_special_token_66|>",
597
- "lstrip": false,
598
- "normalized": false,
599
- "rstrip": false,
600
- "single_word": false,
601
- "special": true
602
- },
603
- "128075": {
604
- "content": "<|reserved_special_token_67|>",
605
- "lstrip": false,
606
- "normalized": false,
607
- "rstrip": false,
608
- "single_word": false,
609
- "special": true
610
- },
611
- "128076": {
612
- "content": "<|reserved_special_token_68|>",
613
- "lstrip": false,
614
- "normalized": false,
615
- "rstrip": false,
616
- "single_word": false,
617
- "special": true
618
- },
619
- "128077": {
620
- "content": "<|reserved_special_token_69|>",
621
- "lstrip": false,
622
- "normalized": false,
623
- "rstrip": false,
624
- "single_word": false,
625
- "special": true
626
- },
627
- "128078": {
628
- "content": "<|reserved_special_token_70|>",
629
- "lstrip": false,
630
- "normalized": false,
631
- "rstrip": false,
632
- "single_word": false,
633
- "special": true
634
- },
635
- "128079": {
636
- "content": "<|reserved_special_token_71|>",
637
- "lstrip": false,
638
- "normalized": false,
639
- "rstrip": false,
640
- "single_word": false,
641
- "special": true
642
- },
643
- "128080": {
644
- "content": "<|reserved_special_token_72|>",
645
- "lstrip": false,
646
- "normalized": false,
647
- "rstrip": false,
648
- "single_word": false,
649
- "special": true
650
- },
651
- "128081": {
652
- "content": "<|reserved_special_token_73|>",
653
- "lstrip": false,
654
- "normalized": false,
655
- "rstrip": false,
656
- "single_word": false,
657
- "special": true
658
- },
659
- "128082": {
660
- "content": "<|reserved_special_token_74|>",
661
- "lstrip": false,
662
- "normalized": false,
663
- "rstrip": false,
664
- "single_word": false,
665
- "special": true
666
- },
667
- "128083": {
668
- "content": "<|reserved_special_token_75|>",
669
- "lstrip": false,
670
- "normalized": false,
671
- "rstrip": false,
672
- "single_word": false,
673
- "special": true
674
- },
675
- "128084": {
676
- "content": "<|reserved_special_token_76|>",
677
- "lstrip": false,
678
- "normalized": false,
679
- "rstrip": false,
680
- "single_word": false,
681
- "special": true
682
- },
683
- "128085": {
684
- "content": "<|reserved_special_token_77|>",
685
- "lstrip": false,
686
- "normalized": false,
687
- "rstrip": false,
688
- "single_word": false,
689
- "special": true
690
- },
691
- "128086": {
692
- "content": "<|reserved_special_token_78|>",
693
- "lstrip": false,
694
- "normalized": false,
695
- "rstrip": false,
696
- "single_word": false,
697
- "special": true
698
- },
699
- "128087": {
700
- "content": "<|reserved_special_token_79|>",
701
- "lstrip": false,
702
- "normalized": false,
703
- "rstrip": false,
704
- "single_word": false,
705
- "special": true
706
- },
707
- "128088": {
708
- "content": "<|reserved_special_token_80|>",
709
- "lstrip": false,
710
- "normalized": false,
711
- "rstrip": false,
712
- "single_word": false,
713
- "special": true
714
- },
715
- "128089": {
716
- "content": "<|reserved_special_token_81|>",
717
- "lstrip": false,
718
- "normalized": false,
719
- "rstrip": false,
720
- "single_word": false,
721
- "special": true
722
- },
723
- "128090": {
724
- "content": "<|reserved_special_token_82|>",
725
- "lstrip": false,
726
- "normalized": false,
727
- "rstrip": false,
728
- "single_word": false,
729
- "special": true
730
- },
731
- "128091": {
732
- "content": "<|reserved_special_token_83|>",
733
- "lstrip": false,
734
- "normalized": false,
735
- "rstrip": false,
736
- "single_word": false,
737
- "special": true
738
- },
739
- "128092": {
740
- "content": "<|reserved_special_token_84|>",
741
- "lstrip": false,
742
- "normalized": false,
743
- "rstrip": false,
744
- "single_word": false,
745
- "special": true
746
- },
747
- "128093": {
748
- "content": "<|reserved_special_token_85|>",
749
- "lstrip": false,
750
- "normalized": false,
751
- "rstrip": false,
752
- "single_word": false,
753
- "special": true
754
- },
755
- "128094": {
756
- "content": "<|reserved_special_token_86|>",
757
- "lstrip": false,
758
- "normalized": false,
759
- "rstrip": false,
760
- "single_word": false,
761
- "special": true
762
- },
763
- "128095": {
764
- "content": "<|reserved_special_token_87|>",
765
- "lstrip": false,
766
- "normalized": false,
767
- "rstrip": false,
768
- "single_word": false,
769
- "special": true
770
- },
771
- "128096": {
772
- "content": "<|reserved_special_token_88|>",
773
- "lstrip": false,
774
- "normalized": false,
775
- "rstrip": false,
776
- "single_word": false,
777
- "special": true
778
- },
779
- "128097": {
780
- "content": "<|reserved_special_token_89|>",
781
- "lstrip": false,
782
- "normalized": false,
783
- "rstrip": false,
784
- "single_word": false,
785
- "special": true
786
- },
787
- "128098": {
788
- "content": "<|reserved_special_token_90|>",
789
- "lstrip": false,
790
- "normalized": false,
791
- "rstrip": false,
792
- "single_word": false,
793
- "special": true
794
- },
795
- "128099": {
796
- "content": "<|reserved_special_token_91|>",
797
- "lstrip": false,
798
- "normalized": false,
799
- "rstrip": false,
800
- "single_word": false,
801
- "special": true
802
- },
803
- "128100": {
804
- "content": "<|reserved_special_token_92|>",
805
- "lstrip": false,
806
- "normalized": false,
807
- "rstrip": false,
808
- "single_word": false,
809
- "special": true
810
- },
811
- "128101": {
812
- "content": "<|reserved_special_token_93|>",
813
- "lstrip": false,
814
- "normalized": false,
815
- "rstrip": false,
816
- "single_word": false,
817
- "special": true
818
- },
819
- "128102": {
820
- "content": "<|reserved_special_token_94|>",
821
- "lstrip": false,
822
- "normalized": false,
823
- "rstrip": false,
824
- "single_word": false,
825
- "special": true
826
- },
827
- "128103": {
828
- "content": "<|reserved_special_token_95|>",
829
- "lstrip": false,
830
- "normalized": false,
831
- "rstrip": false,
832
- "single_word": false,
833
- "special": true
834
- },
835
- "128104": {
836
- "content": "<|reserved_special_token_96|>",
837
- "lstrip": false,
838
- "normalized": false,
839
- "rstrip": false,
840
- "single_word": false,
841
- "special": true
842
- },
843
- "128105": {
844
- "content": "<|reserved_special_token_97|>",
845
- "lstrip": false,
846
- "normalized": false,
847
- "rstrip": false,
848
- "single_word": false,
849
- "special": true
850
- },
851
- "128106": {
852
- "content": "<|reserved_special_token_98|>",
853
- "lstrip": false,
854
- "normalized": false,
855
- "rstrip": false,
856
- "single_word": false,
857
- "special": true
858
- },
859
- "128107": {
860
- "content": "<|reserved_special_token_99|>",
861
- "lstrip": false,
862
- "normalized": false,
863
- "rstrip": false,
864
- "single_word": false,
865
- "special": true
866
- },
867
- "128108": {
868
- "content": "<|reserved_special_token_100|>",
869
- "lstrip": false,
870
- "normalized": false,
871
- "rstrip": false,
872
- "single_word": false,
873
- "special": true
874
- },
875
- "128109": {
876
- "content": "<|reserved_special_token_101|>",
877
- "lstrip": false,
878
- "normalized": false,
879
- "rstrip": false,
880
- "single_word": false,
881
- "special": true
882
- },
883
- "128110": {
884
- "content": "<|reserved_special_token_102|>",
885
- "lstrip": false,
886
- "normalized": false,
887
- "rstrip": false,
888
- "single_word": false,
889
- "special": true
890
- },
891
- "128111": {
892
- "content": "<|reserved_special_token_103|>",
893
- "lstrip": false,
894
- "normalized": false,
895
- "rstrip": false,
896
- "single_word": false,
897
- "special": true
898
- },
899
- "128112": {
900
- "content": "<|reserved_special_token_104|>",
901
- "lstrip": false,
902
- "normalized": false,
903
- "rstrip": false,
904
- "single_word": false,
905
- "special": true
906
- },
907
- "128113": {
908
- "content": "<|reserved_special_token_105|>",
909
- "lstrip": false,
910
- "normalized": false,
911
- "rstrip": false,
912
- "single_word": false,
913
- "special": true
914
- },
915
- "128114": {
916
- "content": "<|reserved_special_token_106|>",
917
- "lstrip": false,
918
- "normalized": false,
919
- "rstrip": false,
920
- "single_word": false,
921
- "special": true
922
- },
923
- "128115": {
924
- "content": "<|reserved_special_token_107|>",
925
- "lstrip": false,
926
- "normalized": false,
927
- "rstrip": false,
928
- "single_word": false,
929
- "special": true
930
- },
931
- "128116": {
932
- "content": "<|reserved_special_token_108|>",
933
- "lstrip": false,
934
- "normalized": false,
935
- "rstrip": false,
936
- "single_word": false,
937
- "special": true
938
- },
939
- "128117": {
940
- "content": "<|reserved_special_token_109|>",
941
- "lstrip": false,
942
- "normalized": false,
943
- "rstrip": false,
944
- "single_word": false,
945
- "special": true
946
- },
947
- "128118": {
948
- "content": "<|reserved_special_token_110|>",
949
- "lstrip": false,
950
- "normalized": false,
951
- "rstrip": false,
952
- "single_word": false,
953
- "special": true
954
- },
955
- "128119": {
956
- "content": "<|reserved_special_token_111|>",
957
- "lstrip": false,
958
- "normalized": false,
959
- "rstrip": false,
960
- "single_word": false,
961
- "special": true
962
- },
963
- "128120": {
964
- "content": "<|reserved_special_token_112|>",
965
- "lstrip": false,
966
- "normalized": false,
967
- "rstrip": false,
968
- "single_word": false,
969
- "special": true
970
- },
971
- "128121": {
972
- "content": "<|reserved_special_token_113|>",
973
- "lstrip": false,
974
- "normalized": false,
975
- "rstrip": false,
976
- "single_word": false,
977
- "special": true
978
- },
979
- "128122": {
980
- "content": "<|reserved_special_token_114|>",
981
- "lstrip": false,
982
- "normalized": false,
983
- "rstrip": false,
984
- "single_word": false,
985
- "special": true
986
- },
987
- "128123": {
988
- "content": "<|reserved_special_token_115|>",
989
- "lstrip": false,
990
- "normalized": false,
991
- "rstrip": false,
992
- "single_word": false,
993
- "special": true
994
- },
995
- "128124": {
996
- "content": "<|reserved_special_token_116|>",
997
- "lstrip": false,
998
- "normalized": false,
999
- "rstrip": false,
1000
- "single_word": false,
1001
- "special": true
1002
- },
1003
- "128125": {
1004
- "content": "<|reserved_special_token_117|>",
1005
- "lstrip": false,
1006
- "normalized": false,
1007
- "rstrip": false,
1008
- "single_word": false,
1009
- "special": true
1010
- },
1011
- "128126": {
1012
- "content": "<|reserved_special_token_118|>",
1013
- "lstrip": false,
1014
- "normalized": false,
1015
- "rstrip": false,
1016
- "single_word": false,
1017
- "special": true
1018
- },
1019
- "128127": {
1020
- "content": "<|reserved_special_token_119|>",
1021
- "lstrip": false,
1022
- "normalized": false,
1023
- "rstrip": false,
1024
- "single_word": false,
1025
- "special": true
1026
- },
1027
- "128128": {
1028
- "content": "<|reserved_special_token_120|>",
1029
- "lstrip": false,
1030
- "normalized": false,
1031
- "rstrip": false,
1032
- "single_word": false,
1033
- "special": true
1034
- },
1035
- "128129": {
1036
- "content": "<|reserved_special_token_121|>",
1037
- "lstrip": false,
1038
- "normalized": false,
1039
- "rstrip": false,
1040
- "single_word": false,
1041
- "special": true
1042
- },
1043
- "128130": {
1044
- "content": "<|reserved_special_token_122|>",
1045
- "lstrip": false,
1046
- "normalized": false,
1047
- "rstrip": false,
1048
- "single_word": false,
1049
- "special": true
1050
- },
1051
- "128131": {
1052
- "content": "<|reserved_special_token_123|>",
1053
- "lstrip": false,
1054
- "normalized": false,
1055
- "rstrip": false,
1056
- "single_word": false,
1057
- "special": true
1058
- },
1059
- "128132": {
1060
- "content": "<|reserved_special_token_124|>",
1061
- "lstrip": false,
1062
- "normalized": false,
1063
- "rstrip": false,
1064
- "single_word": false,
1065
- "special": true
1066
- },
1067
- "128133": {
1068
- "content": "<|reserved_special_token_125|>",
1069
- "lstrip": false,
1070
- "normalized": false,
1071
- "rstrip": false,
1072
- "single_word": false,
1073
- "special": true
1074
- },
1075
- "128134": {
1076
- "content": "<|reserved_special_token_126|>",
1077
- "lstrip": false,
1078
- "normalized": false,
1079
- "rstrip": false,
1080
- "single_word": false,
1081
- "special": true
1082
- },
1083
- "128135": {
1084
- "content": "<|reserved_special_token_127|>",
1085
- "lstrip": false,
1086
- "normalized": false,
1087
- "rstrip": false,
1088
- "single_word": false,
1089
- "special": true
1090
- },
1091
- "128136": {
1092
- "content": "<|reserved_special_token_128|>",
1093
- "lstrip": false,
1094
- "normalized": false,
1095
- "rstrip": false,
1096
- "single_word": false,
1097
- "special": true
1098
- },
1099
- "128137": {
1100
- "content": "<|reserved_special_token_129|>",
1101
- "lstrip": false,
1102
- "normalized": false,
1103
- "rstrip": false,
1104
- "single_word": false,
1105
- "special": true
1106
- },
1107
- "128138": {
1108
- "content": "<|reserved_special_token_130|>",
1109
- "lstrip": false,
1110
- "normalized": false,
1111
- "rstrip": false,
1112
- "single_word": false,
1113
- "special": true
1114
- },
1115
- "128139": {
1116
- "content": "<|reserved_special_token_131|>",
1117
- "lstrip": false,
1118
- "normalized": false,
1119
- "rstrip": false,
1120
- "single_word": false,
1121
- "special": true
1122
- },
1123
- "128140": {
1124
- "content": "<|reserved_special_token_132|>",
1125
- "lstrip": false,
1126
- "normalized": false,
1127
- "rstrip": false,
1128
- "single_word": false,
1129
- "special": true
1130
- },
1131
- "128141": {
1132
- "content": "<|reserved_special_token_133|>",
1133
- "lstrip": false,
1134
- "normalized": false,
1135
- "rstrip": false,
1136
- "single_word": false,
1137
- "special": true
1138
- },
1139
- "128142": {
1140
- "content": "<|reserved_special_token_134|>",
1141
- "lstrip": false,
1142
- "normalized": false,
1143
- "rstrip": false,
1144
- "single_word": false,
1145
- "special": true
1146
- },
1147
- "128143": {
1148
- "content": "<|reserved_special_token_135|>",
1149
- "lstrip": false,
1150
- "normalized": false,
1151
- "rstrip": false,
1152
- "single_word": false,
1153
- "special": true
1154
- },
1155
- "128144": {
1156
- "content": "<|reserved_special_token_136|>",
1157
- "lstrip": false,
1158
- "normalized": false,
1159
- "rstrip": false,
1160
- "single_word": false,
1161
- "special": true
1162
- },
1163
- "128145": {
1164
- "content": "<|reserved_special_token_137|>",
1165
- "lstrip": false,
1166
- "normalized": false,
1167
- "rstrip": false,
1168
- "single_word": false,
1169
- "special": true
1170
- },
1171
- "128146": {
1172
- "content": "<|reserved_special_token_138|>",
1173
- "lstrip": false,
1174
- "normalized": false,
1175
- "rstrip": false,
1176
- "single_word": false,
1177
- "special": true
1178
- },
1179
- "128147": {
1180
- "content": "<|reserved_special_token_139|>",
1181
- "lstrip": false,
1182
- "normalized": false,
1183
- "rstrip": false,
1184
- "single_word": false,
1185
- "special": true
1186
- },
1187
- "128148": {
1188
- "content": "<|reserved_special_token_140|>",
1189
- "lstrip": false,
1190
- "normalized": false,
1191
- "rstrip": false,
1192
- "single_word": false,
1193
- "special": true
1194
- },
1195
- "128149": {
1196
- "content": "<|reserved_special_token_141|>",
1197
- "lstrip": false,
1198
- "normalized": false,
1199
- "rstrip": false,
1200
- "single_word": false,
1201
- "special": true
1202
- },
1203
- "128150": {
1204
- "content": "<|reserved_special_token_142|>",
1205
- "lstrip": false,
1206
- "normalized": false,
1207
- "rstrip": false,
1208
- "single_word": false,
1209
- "special": true
1210
- },
1211
- "128151": {
1212
- "content": "<|reserved_special_token_143|>",
1213
- "lstrip": false,
1214
- "normalized": false,
1215
- "rstrip": false,
1216
- "single_word": false,
1217
- "special": true
1218
- },
1219
- "128152": {
1220
- "content": "<|reserved_special_token_144|>",
1221
- "lstrip": false,
1222
- "normalized": false,
1223
- "rstrip": false,
1224
- "single_word": false,
1225
- "special": true
1226
- },
1227
- "128153": {
1228
- "content": "<|reserved_special_token_145|>",
1229
- "lstrip": false,
1230
- "normalized": false,
1231
- "rstrip": false,
1232
- "single_word": false,
1233
- "special": true
1234
- },
1235
- "128154": {
1236
- "content": "<|reserved_special_token_146|>",
1237
- "lstrip": false,
1238
- "normalized": false,
1239
- "rstrip": false,
1240
- "single_word": false,
1241
- "special": true
1242
- },
1243
- "128155": {
1244
- "content": "<|reserved_special_token_147|>",
1245
- "lstrip": false,
1246
- "normalized": false,
1247
- "rstrip": false,
1248
- "single_word": false,
1249
- "special": true
1250
- },
1251
- "128156": {
1252
- "content": "<|reserved_special_token_148|>",
1253
- "lstrip": false,
1254
- "normalized": false,
1255
- "rstrip": false,
1256
- "single_word": false,
1257
- "special": true
1258
- },
1259
- "128157": {
1260
- "content": "<|reserved_special_token_149|>",
1261
- "lstrip": false,
1262
- "normalized": false,
1263
- "rstrip": false,
1264
- "single_word": false,
1265
- "special": true
1266
- },
1267
- "128158": {
1268
- "content": "<|reserved_special_token_150|>",
1269
- "lstrip": false,
1270
- "normalized": false,
1271
- "rstrip": false,
1272
- "single_word": false,
1273
- "special": true
1274
- },
1275
- "128159": {
1276
- "content": "<|reserved_special_token_151|>",
1277
- "lstrip": false,
1278
- "normalized": false,
1279
- "rstrip": false,
1280
- "single_word": false,
1281
- "special": true
1282
- },
1283
- "128160": {
1284
- "content": "<|reserved_special_token_152|>",
1285
- "lstrip": false,
1286
- "normalized": false,
1287
- "rstrip": false,
1288
- "single_word": false,
1289
- "special": true
1290
- },
1291
- "128161": {
1292
- "content": "<|reserved_special_token_153|>",
1293
- "lstrip": false,
1294
- "normalized": false,
1295
- "rstrip": false,
1296
- "single_word": false,
1297
- "special": true
1298
- },
1299
- "128162": {
1300
- "content": "<|reserved_special_token_154|>",
1301
- "lstrip": false,
1302
- "normalized": false,
1303
- "rstrip": false,
1304
- "single_word": false,
1305
- "special": true
1306
- },
1307
- "128163": {
1308
- "content": "<|reserved_special_token_155|>",
1309
- "lstrip": false,
1310
- "normalized": false,
1311
- "rstrip": false,
1312
- "single_word": false,
1313
- "special": true
1314
- },
1315
- "128164": {
1316
- "content": "<|reserved_special_token_156|>",
1317
- "lstrip": false,
1318
- "normalized": false,
1319
- "rstrip": false,
1320
- "single_word": false,
1321
- "special": true
1322
- },
1323
- "128165": {
1324
- "content": "<|reserved_special_token_157|>",
1325
- "lstrip": false,
1326
- "normalized": false,
1327
- "rstrip": false,
1328
- "single_word": false,
1329
- "special": true
1330
- },
1331
- "128166": {
1332
- "content": "<|reserved_special_token_158|>",
1333
- "lstrip": false,
1334
- "normalized": false,
1335
- "rstrip": false,
1336
- "single_word": false,
1337
- "special": true
1338
- },
1339
- "128167": {
1340
- "content": "<|reserved_special_token_159|>",
1341
- "lstrip": false,
1342
- "normalized": false,
1343
- "rstrip": false,
1344
- "single_word": false,
1345
- "special": true
1346
- },
1347
- "128168": {
1348
- "content": "<|reserved_special_token_160|>",
1349
- "lstrip": false,
1350
- "normalized": false,
1351
- "rstrip": false,
1352
- "single_word": false,
1353
- "special": true
1354
- },
1355
- "128169": {
1356
- "content": "<|reserved_special_token_161|>",
1357
- "lstrip": false,
1358
- "normalized": false,
1359
- "rstrip": false,
1360
- "single_word": false,
1361
- "special": true
1362
- },
1363
- "128170": {
1364
- "content": "<|reserved_special_token_162|>",
1365
- "lstrip": false,
1366
- "normalized": false,
1367
- "rstrip": false,
1368
- "single_word": false,
1369
- "special": true
1370
- },
1371
- "128171": {
1372
- "content": "<|reserved_special_token_163|>",
1373
- "lstrip": false,
1374
- "normalized": false,
1375
- "rstrip": false,
1376
- "single_word": false,
1377
- "special": true
1378
- },
1379
- "128172": {
1380
- "content": "<|reserved_special_token_164|>",
1381
- "lstrip": false,
1382
- "normalized": false,
1383
- "rstrip": false,
1384
- "single_word": false,
1385
- "special": true
1386
- },
1387
- "128173": {
1388
- "content": "<|reserved_special_token_165|>",
1389
- "lstrip": false,
1390
- "normalized": false,
1391
- "rstrip": false,
1392
- "single_word": false,
1393
- "special": true
1394
- },
1395
- "128174": {
1396
- "content": "<|reserved_special_token_166|>",
1397
- "lstrip": false,
1398
- "normalized": false,
1399
- "rstrip": false,
1400
- "single_word": false,
1401
- "special": true
1402
- },
1403
- "128175": {
1404
- "content": "<|reserved_special_token_167|>",
1405
- "lstrip": false,
1406
- "normalized": false,
1407
- "rstrip": false,
1408
- "single_word": false,
1409
- "special": true
1410
- },
1411
- "128176": {
1412
- "content": "<|reserved_special_token_168|>",
1413
- "lstrip": false,
1414
- "normalized": false,
1415
- "rstrip": false,
1416
- "single_word": false,
1417
- "special": true
1418
- },
1419
- "128177": {
1420
- "content": "<|reserved_special_token_169|>",
1421
- "lstrip": false,
1422
- "normalized": false,
1423
- "rstrip": false,
1424
- "single_word": false,
1425
- "special": true
1426
- },
1427
- "128178": {
1428
- "content": "<|reserved_special_token_170|>",
1429
- "lstrip": false,
1430
- "normalized": false,
1431
- "rstrip": false,
1432
- "single_word": false,
1433
- "special": true
1434
- },
1435
- "128179": {
1436
- "content": "<|reserved_special_token_171|>",
1437
- "lstrip": false,
1438
- "normalized": false,
1439
- "rstrip": false,
1440
- "single_word": false,
1441
- "special": true
1442
- },
1443
- "128180": {
1444
- "content": "<|reserved_special_token_172|>",
1445
- "lstrip": false,
1446
- "normalized": false,
1447
- "rstrip": false,
1448
- "single_word": false,
1449
- "special": true
1450
- },
1451
- "128181": {
1452
- "content": "<|reserved_special_token_173|>",
1453
- "lstrip": false,
1454
- "normalized": false,
1455
- "rstrip": false,
1456
- "single_word": false,
1457
- "special": true
1458
- },
1459
- "128182": {
1460
- "content": "<|reserved_special_token_174|>",
1461
- "lstrip": false,
1462
- "normalized": false,
1463
- "rstrip": false,
1464
- "single_word": false,
1465
- "special": true
1466
- },
1467
- "128183": {
1468
- "content": "<|reserved_special_token_175|>",
1469
- "lstrip": false,
1470
- "normalized": false,
1471
- "rstrip": false,
1472
- "single_word": false,
1473
- "special": true
1474
- },
1475
- "128184": {
1476
- "content": "<|reserved_special_token_176|>",
1477
- "lstrip": false,
1478
- "normalized": false,
1479
- "rstrip": false,
1480
- "single_word": false,
1481
- "special": true
1482
- },
1483
- "128185": {
1484
- "content": "<|reserved_special_token_177|>",
1485
- "lstrip": false,
1486
- "normalized": false,
1487
- "rstrip": false,
1488
- "single_word": false,
1489
- "special": true
1490
- },
1491
- "128186": {
1492
- "content": "<|reserved_special_token_178|>",
1493
- "lstrip": false,
1494
- "normalized": false,
1495
- "rstrip": false,
1496
- "single_word": false,
1497
- "special": true
1498
- },
1499
- "128187": {
1500
- "content": "<|reserved_special_token_179|>",
1501
- "lstrip": false,
1502
- "normalized": false,
1503
- "rstrip": false,
1504
- "single_word": false,
1505
- "special": true
1506
- },
1507
- "128188": {
1508
- "content": "<|reserved_special_token_180|>",
1509
- "lstrip": false,
1510
- "normalized": false,
1511
- "rstrip": false,
1512
- "single_word": false,
1513
- "special": true
1514
- },
1515
- "128189": {
1516
- "content": "<|reserved_special_token_181|>",
1517
- "lstrip": false,
1518
- "normalized": false,
1519
- "rstrip": false,
1520
- "single_word": false,
1521
- "special": true
1522
- },
1523
- "128190": {
1524
- "content": "<|reserved_special_token_182|>",
1525
- "lstrip": false,
1526
- "normalized": false,
1527
- "rstrip": false,
1528
- "single_word": false,
1529
- "special": true
1530
- },
1531
- "128191": {
1532
- "content": "<|reserved_special_token_183|>",
1533
- "lstrip": false,
1534
- "normalized": false,
1535
- "rstrip": false,
1536
- "single_word": false,
1537
- "special": true
1538
- },
1539
- "128192": {
1540
- "content": "<|reserved_special_token_184|>",
1541
- "lstrip": false,
1542
- "normalized": false,
1543
- "rstrip": false,
1544
- "single_word": false,
1545
- "special": true
1546
- },
1547
- "128193": {
1548
- "content": "<|reserved_special_token_185|>",
1549
- "lstrip": false,
1550
- "normalized": false,
1551
- "rstrip": false,
1552
- "single_word": false,
1553
- "special": true
1554
- },
1555
- "128194": {
1556
- "content": "<|reserved_special_token_186|>",
1557
- "lstrip": false,
1558
- "normalized": false,
1559
- "rstrip": false,
1560
- "single_word": false,
1561
- "special": true
1562
- },
1563
- "128195": {
1564
- "content": "<|reserved_special_token_187|>",
1565
- "lstrip": false,
1566
- "normalized": false,
1567
- "rstrip": false,
1568
- "single_word": false,
1569
- "special": true
1570
- },
1571
- "128196": {
1572
- "content": "<|reserved_special_token_188|>",
1573
- "lstrip": false,
1574
- "normalized": false,
1575
- "rstrip": false,
1576
- "single_word": false,
1577
- "special": true
1578
- },
1579
- "128197": {
1580
- "content": "<|reserved_special_token_189|>",
1581
- "lstrip": false,
1582
- "normalized": false,
1583
- "rstrip": false,
1584
- "single_word": false,
1585
- "special": true
1586
- },
1587
- "128198": {
1588
- "content": "<|reserved_special_token_190|>",
1589
- "lstrip": false,
1590
- "normalized": false,
1591
- "rstrip": false,
1592
- "single_word": false,
1593
- "special": true
1594
- },
1595
- "128199": {
1596
- "content": "<|reserved_special_token_191|>",
1597
- "lstrip": false,
1598
- "normalized": false,
1599
- "rstrip": false,
1600
- "single_word": false,
1601
- "special": true
1602
- },
1603
- "128200": {
1604
- "content": "<|reserved_special_token_192|>",
1605
- "lstrip": false,
1606
- "normalized": false,
1607
- "rstrip": false,
1608
- "single_word": false,
1609
- "special": true
1610
- },
1611
- "128201": {
1612
- "content": "<|reserved_special_token_193|>",
1613
- "lstrip": false,
1614
- "normalized": false,
1615
- "rstrip": false,
1616
- "single_word": false,
1617
- "special": true
1618
- },
1619
- "128202": {
1620
- "content": "<|reserved_special_token_194|>",
1621
- "lstrip": false,
1622
- "normalized": false,
1623
- "rstrip": false,
1624
- "single_word": false,
1625
- "special": true
1626
- },
1627
- "128203": {
1628
- "content": "<|reserved_special_token_195|>",
1629
- "lstrip": false,
1630
- "normalized": false,
1631
- "rstrip": false,
1632
- "single_word": false,
1633
- "special": true
1634
- },
1635
- "128204": {
1636
- "content": "<|reserved_special_token_196|>",
1637
- "lstrip": false,
1638
- "normalized": false,
1639
- "rstrip": false,
1640
- "single_word": false,
1641
- "special": true
1642
- },
1643
- "128205": {
1644
- "content": "<|reserved_special_token_197|>",
1645
- "lstrip": false,
1646
- "normalized": false,
1647
- "rstrip": false,
1648
- "single_word": false,
1649
- "special": true
1650
- },
1651
- "128206": {
1652
- "content": "<|reserved_special_token_198|>",
1653
- "lstrip": false,
1654
- "normalized": false,
1655
- "rstrip": false,
1656
- "single_word": false,
1657
- "special": true
1658
- },
1659
- "128207": {
1660
- "content": "<|reserved_special_token_199|>",
1661
- "lstrip": false,
1662
- "normalized": false,
1663
- "rstrip": false,
1664
- "single_word": false,
1665
- "special": true
1666
- },
1667
- "128208": {
1668
- "content": "<|reserved_special_token_200|>",
1669
- "lstrip": false,
1670
- "normalized": false,
1671
- "rstrip": false,
1672
- "single_word": false,
1673
- "special": true
1674
- },
1675
- "128209": {
1676
- "content": "<|reserved_special_token_201|>",
1677
- "lstrip": false,
1678
- "normalized": false,
1679
- "rstrip": false,
1680
- "single_word": false,
1681
- "special": true
1682
- },
1683
- "128210": {
1684
- "content": "<|reserved_special_token_202|>",
1685
- "lstrip": false,
1686
- "normalized": false,
1687
- "rstrip": false,
1688
- "single_word": false,
1689
- "special": true
1690
- },
1691
- "128211": {
1692
- "content": "<|reserved_special_token_203|>",
1693
- "lstrip": false,
1694
- "normalized": false,
1695
- "rstrip": false,
1696
- "single_word": false,
1697
- "special": true
1698
- },
1699
- "128212": {
1700
- "content": "<|reserved_special_token_204|>",
1701
- "lstrip": false,
1702
- "normalized": false,
1703
- "rstrip": false,
1704
- "single_word": false,
1705
- "special": true
1706
- },
1707
- "128213": {
1708
- "content": "<|reserved_special_token_205|>",
1709
- "lstrip": false,
1710
- "normalized": false,
1711
- "rstrip": false,
1712
- "single_word": false,
1713
- "special": true
1714
- },
1715
- "128214": {
1716
- "content": "<|reserved_special_token_206|>",
1717
- "lstrip": false,
1718
- "normalized": false,
1719
- "rstrip": false,
1720
- "single_word": false,
1721
- "special": true
1722
- },
1723
- "128215": {
1724
- "content": "<|reserved_special_token_207|>",
1725
- "lstrip": false,
1726
- "normalized": false,
1727
- "rstrip": false,
1728
- "single_word": false,
1729
- "special": true
1730
- },
1731
- "128216": {
1732
- "content": "<|reserved_special_token_208|>",
1733
- "lstrip": false,
1734
- "normalized": false,
1735
- "rstrip": false,
1736
- "single_word": false,
1737
- "special": true
1738
- },
1739
- "128217": {
1740
- "content": "<|reserved_special_token_209|>",
1741
- "lstrip": false,
1742
- "normalized": false,
1743
- "rstrip": false,
1744
- "single_word": false,
1745
- "special": true
1746
- },
1747
- "128218": {
1748
- "content": "<|reserved_special_token_210|>",
1749
- "lstrip": false,
1750
- "normalized": false,
1751
- "rstrip": false,
1752
- "single_word": false,
1753
- "special": true
1754
- },
1755
- "128219": {
1756
- "content": "<|reserved_special_token_211|>",
1757
- "lstrip": false,
1758
- "normalized": false,
1759
- "rstrip": false,
1760
- "single_word": false,
1761
- "special": true
1762
- },
1763
- "128220": {
1764
- "content": "<|reserved_special_token_212|>",
1765
- "lstrip": false,
1766
- "normalized": false,
1767
- "rstrip": false,
1768
- "single_word": false,
1769
- "special": true
1770
- },
1771
- "128221": {
1772
- "content": "<|reserved_special_token_213|>",
1773
- "lstrip": false,
1774
- "normalized": false,
1775
- "rstrip": false,
1776
- "single_word": false,
1777
- "special": true
1778
- },
1779
- "128222": {
1780
- "content": "<|reserved_special_token_214|>",
1781
- "lstrip": false,
1782
- "normalized": false,
1783
- "rstrip": false,
1784
- "single_word": false,
1785
- "special": true
1786
- },
1787
- "128223": {
1788
- "content": "<|reserved_special_token_215|>",
1789
- "lstrip": false,
1790
- "normalized": false,
1791
- "rstrip": false,
1792
- "single_word": false,
1793
- "special": true
1794
- },
1795
- "128224": {
1796
- "content": "<|reserved_special_token_216|>",
1797
- "lstrip": false,
1798
- "normalized": false,
1799
- "rstrip": false,
1800
- "single_word": false,
1801
- "special": true
1802
- },
1803
- "128225": {
1804
- "content": "<|reserved_special_token_217|>",
1805
- "lstrip": false,
1806
- "normalized": false,
1807
- "rstrip": false,
1808
- "single_word": false,
1809
- "special": true
1810
- },
1811
- "128226": {
1812
- "content": "<|reserved_special_token_218|>",
1813
- "lstrip": false,
1814
- "normalized": false,
1815
- "rstrip": false,
1816
- "single_word": false,
1817
- "special": true
1818
- },
1819
- "128227": {
1820
- "content": "<|reserved_special_token_219|>",
1821
- "lstrip": false,
1822
- "normalized": false,
1823
- "rstrip": false,
1824
- "single_word": false,
1825
- "special": true
1826
- },
1827
- "128228": {
1828
- "content": "<|reserved_special_token_220|>",
1829
- "lstrip": false,
1830
- "normalized": false,
1831
- "rstrip": false,
1832
- "single_word": false,
1833
- "special": true
1834
- },
1835
- "128229": {
1836
- "content": "<|reserved_special_token_221|>",
1837
- "lstrip": false,
1838
- "normalized": false,
1839
- "rstrip": false,
1840
- "single_word": false,
1841
- "special": true
1842
- },
1843
- "128230": {
1844
- "content": "<|reserved_special_token_222|>",
1845
- "lstrip": false,
1846
- "normalized": false,
1847
- "rstrip": false,
1848
- "single_word": false,
1849
- "special": true
1850
- },
1851
- "128231": {
1852
- "content": "<|reserved_special_token_223|>",
1853
- "lstrip": false,
1854
- "normalized": false,
1855
- "rstrip": false,
1856
- "single_word": false,
1857
- "special": true
1858
- },
1859
- "128232": {
1860
- "content": "<|reserved_special_token_224|>",
1861
- "lstrip": false,
1862
- "normalized": false,
1863
- "rstrip": false,
1864
- "single_word": false,
1865
- "special": true
1866
- },
1867
- "128233": {
1868
- "content": "<|reserved_special_token_225|>",
1869
- "lstrip": false,
1870
- "normalized": false,
1871
- "rstrip": false,
1872
- "single_word": false,
1873
- "special": true
1874
- },
1875
- "128234": {
1876
- "content": "<|reserved_special_token_226|>",
1877
- "lstrip": false,
1878
- "normalized": false,
1879
- "rstrip": false,
1880
- "single_word": false,
1881
- "special": true
1882
- },
1883
- "128235": {
1884
- "content": "<|reserved_special_token_227|>",
1885
- "lstrip": false,
1886
- "normalized": false,
1887
- "rstrip": false,
1888
- "single_word": false,
1889
- "special": true
1890
- },
1891
- "128236": {
1892
- "content": "<|reserved_special_token_228|>",
1893
- "lstrip": false,
1894
- "normalized": false,
1895
- "rstrip": false,
1896
- "single_word": false,
1897
- "special": true
1898
- },
1899
- "128237": {
1900
- "content": "<|reserved_special_token_229|>",
1901
- "lstrip": false,
1902
- "normalized": false,
1903
- "rstrip": false,
1904
- "single_word": false,
1905
- "special": true
1906
- },
1907
- "128238": {
1908
- "content": "<|reserved_special_token_230|>",
1909
- "lstrip": false,
1910
- "normalized": false,
1911
- "rstrip": false,
1912
- "single_word": false,
1913
- "special": true
1914
- },
1915
- "128239": {
1916
- "content": "<|reserved_special_token_231|>",
1917
- "lstrip": false,
1918
- "normalized": false,
1919
- "rstrip": false,
1920
- "single_word": false,
1921
- "special": true
1922
- },
1923
- "128240": {
1924
- "content": "<|reserved_special_token_232|>",
1925
- "lstrip": false,
1926
- "normalized": false,
1927
- "rstrip": false,
1928
- "single_word": false,
1929
- "special": true
1930
- },
1931
- "128241": {
1932
- "content": "<|reserved_special_token_233|>",
1933
- "lstrip": false,
1934
- "normalized": false,
1935
- "rstrip": false,
1936
- "single_word": false,
1937
- "special": true
1938
- },
1939
- "128242": {
1940
- "content": "<|reserved_special_token_234|>",
1941
- "lstrip": false,
1942
- "normalized": false,
1943
- "rstrip": false,
1944
- "single_word": false,
1945
- "special": true
1946
- },
1947
- "128243": {
1948
- "content": "<|reserved_special_token_235|>",
1949
- "lstrip": false,
1950
- "normalized": false,
1951
- "rstrip": false,
1952
- "single_word": false,
1953
- "special": true
1954
- },
1955
- "128244": {
1956
- "content": "<|reserved_special_token_236|>",
1957
- "lstrip": false,
1958
- "normalized": false,
1959
- "rstrip": false,
1960
- "single_word": false,
1961
- "special": true
1962
- },
1963
- "128245": {
1964
- "content": "<|reserved_special_token_237|>",
1965
- "lstrip": false,
1966
- "normalized": false,
1967
- "rstrip": false,
1968
- "single_word": false,
1969
- "special": true
1970
- },
1971
- "128246": {
1972
- "content": "<|reserved_special_token_238|>",
1973
- "lstrip": false,
1974
- "normalized": false,
1975
- "rstrip": false,
1976
- "single_word": false,
1977
- "special": true
1978
- },
1979
- "128247": {
1980
- "content": "<|reserved_special_token_239|>",
1981
- "lstrip": false,
1982
- "normalized": false,
1983
- "rstrip": false,
1984
- "single_word": false,
1985
- "special": true
1986
- },
1987
- "128248": {
1988
- "content": "<|reserved_special_token_240|>",
1989
- "lstrip": false,
1990
- "normalized": false,
1991
- "rstrip": false,
1992
- "single_word": false,
1993
- "special": true
1994
- },
1995
- "128249": {
1996
- "content": "<|reserved_special_token_241|>",
1997
- "lstrip": false,
1998
- "normalized": false,
1999
- "rstrip": false,
2000
- "single_word": false,
2001
- "special": true
2002
- },
2003
- "128250": {
2004
- "content": "<|reserved_special_token_242|>",
2005
- "lstrip": false,
2006
- "normalized": false,
2007
- "rstrip": false,
2008
- "single_word": false,
2009
- "special": true
2010
- },
2011
- "128251": {
2012
- "content": "<|reserved_special_token_243|>",
2013
- "lstrip": false,
2014
- "normalized": false,
2015
- "rstrip": false,
2016
- "single_word": false,
2017
- "special": true
2018
- },
2019
- "128252": {
2020
- "content": "<|reserved_special_token_244|>",
2021
- "lstrip": false,
2022
- "normalized": false,
2023
- "rstrip": false,
2024
- "single_word": false,
2025
- "special": true
2026
- },
2027
- "128253": {
2028
- "content": "<|reserved_special_token_245|>",
2029
- "lstrip": false,
2030
- "normalized": false,
2031
- "rstrip": false,
2032
- "single_word": false,
2033
- "special": true
2034
- },
2035
- "128254": {
2036
- "content": "<|reserved_special_token_246|>",
2037
- "lstrip": false,
2038
- "normalized": false,
2039
- "rstrip": false,
2040
- "single_word": false,
2041
- "special": true
2042
- },
2043
- "128255": {
2044
- "content": "<|reserved_special_token_247|>",
2045
- "lstrip": false,
2046
- "normalized": false,
2047
- "rstrip": false,
2048
- "single_word": false,
2049
- "special": true
2050
- }
2051
- },
2052
- "additional_special_tokens": [
2053
- "<|im_start|>",
2054
- "<|im_end|>"
2055
- ],
2056
- "bos_token": "<|im_start|>",
2057
- "chat_template": "{% for message in messages %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}",
2058
- "clean_up_tokenization_spaces": true,
2059
- "eos_token": "<|im_end|>",
2060
- "model_input_names": [
2061
- "input_ids",
2062
- "attention_mask"
2063
- ],
2064
- "model_max_length": 131072,
2065
- "pad_token": "<|im_end|>",
2066
- "tokenizer_class": "PreTrainedTokenizerFast"
2067
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
checkpoint-250/trainer_state.json DELETED
@@ -1,1823 +0,0 @@
1
- {
2
- "best_metric": null,
3
- "best_model_checkpoint": null,
4
- "epoch": 5.0,
5
- "eval_steps": 500,
6
- "global_step": 250,
7
- "is_hyper_param_search": false,
8
- "is_local_process_zero": true,
9
- "is_world_process_zero": true,
10
- "log_history": [
11
- {
12
- "epoch": 0.02,
13
- "grad_norm": 2.7934532165527344,
14
- "learning_rate": 0.0002,
15
- "loss": 2.3404,
16
- "step": 1
17
- },
18
- {
19
- "epoch": 0.04,
20
- "grad_norm": 1.5168005228042603,
21
- "learning_rate": 0.0002,
22
- "loss": 2.0804,
23
- "step": 2
24
- },
25
- {
26
- "epoch": 0.06,
27
- "grad_norm": 1.047807216644287,
28
- "learning_rate": 0.0002,
29
- "loss": 1.9184,
30
- "step": 3
31
- },
32
- {
33
- "epoch": 0.08,
34
- "grad_norm": 1.041599988937378,
35
- "learning_rate": 0.0002,
36
- "loss": 2.0393,
37
- "step": 4
38
- },
39
- {
40
- "epoch": 0.1,
41
- "grad_norm": 0.8074644804000854,
42
- "learning_rate": 0.0002,
43
- "loss": 2.1779,
44
- "step": 5
45
- },
46
- {
47
- "epoch": 0.12,
48
- "grad_norm": 0.7784727811813354,
49
- "learning_rate": 0.0002,
50
- "loss": 2.1583,
51
- "step": 6
52
- },
53
- {
54
- "epoch": 0.14,
55
- "grad_norm": 0.8535248637199402,
56
- "learning_rate": 0.0002,
57
- "loss": 2.0766,
58
- "step": 7
59
- },
60
- {
61
- "epoch": 0.16,
62
- "grad_norm": 0.8507911562919617,
63
- "learning_rate": 0.0002,
64
- "loss": 2.1175,
65
- "step": 8
66
- },
67
- {
68
- "epoch": 0.18,
69
- "grad_norm": 0.8497746586799622,
70
- "learning_rate": 0.0002,
71
- "loss": 2.1156,
72
- "step": 9
73
- },
74
- {
75
- "epoch": 0.2,
76
- "grad_norm": 0.9205197691917419,
77
- "learning_rate": 0.0002,
78
- "loss": 2.1782,
79
- "step": 10
80
- },
81
- {
82
- "epoch": 0.22,
83
- "grad_norm": 1.2332911491394043,
84
- "learning_rate": 0.0002,
85
- "loss": 2.0536,
86
- "step": 11
87
- },
88
- {
89
- "epoch": 0.24,
90
- "grad_norm": 1.3457396030426025,
91
- "learning_rate": 0.0002,
92
- "loss": 2.2186,
93
- "step": 12
94
- },
95
- {
96
- "epoch": 0.26,
97
- "grad_norm": 0.6494730114936829,
98
- "learning_rate": 0.0002,
99
- "loss": 1.8011,
100
- "step": 13
101
- },
102
- {
103
- "epoch": 0.28,
104
- "grad_norm": 0.6363134980201721,
105
- "learning_rate": 0.0002,
106
- "loss": 1.8819,
107
- "step": 14
108
- },
109
- {
110
- "epoch": 0.3,
111
- "grad_norm": 0.7927612662315369,
112
- "learning_rate": 0.0002,
113
- "loss": 1.8845,
114
- "step": 15
115
- },
116
- {
117
- "epoch": 0.32,
118
- "grad_norm": 0.7082176804542542,
119
- "learning_rate": 0.0002,
120
- "loss": 1.8748,
121
- "step": 16
122
- },
123
- {
124
- "epoch": 0.34,
125
- "grad_norm": 0.861709713935852,
126
- "learning_rate": 0.0002,
127
- "loss": 1.8777,
128
- "step": 17
129
- },
130
- {
131
- "epoch": 0.36,
132
- "grad_norm": 0.7901681661605835,
133
- "learning_rate": 0.0002,
134
- "loss": 1.8811,
135
- "step": 18
136
- },
137
- {
138
- "epoch": 0.38,
139
- "grad_norm": 0.7719288468360901,
140
- "learning_rate": 0.0002,
141
- "loss": 1.8898,
142
- "step": 19
143
- },
144
- {
145
- "epoch": 0.4,
146
- "grad_norm": 1.027469277381897,
147
- "learning_rate": 0.0002,
148
- "loss": 2.0089,
149
- "step": 20
150
- },
151
- {
152
- "epoch": 0.42,
153
- "grad_norm": 0.9486727714538574,
154
- "learning_rate": 0.0002,
155
- "loss": 1.968,
156
- "step": 21
157
- },
158
- {
159
- "epoch": 0.44,
160
- "grad_norm": 0.9629890322685242,
161
- "learning_rate": 0.0002,
162
- "loss": 2.0568,
163
- "step": 22
164
- },
165
- {
166
- "epoch": 0.46,
167
- "grad_norm": 1.033793568611145,
168
- "learning_rate": 0.0002,
169
- "loss": 1.8401,
170
- "step": 23
171
- },
172
- {
173
- "epoch": 0.48,
174
- "grad_norm": 1.3298218250274658,
175
- "learning_rate": 0.0002,
176
- "loss": 1.9165,
177
- "step": 24
178
- },
179
- {
180
- "epoch": 0.5,
181
- "grad_norm": 0.6936089396476746,
182
- "learning_rate": 0.0002,
183
- "loss": 1.6614,
184
- "step": 25
185
- },
186
- {
187
- "epoch": 0.52,
188
- "grad_norm": 0.6136096119880676,
189
- "learning_rate": 0.0002,
190
- "loss": 1.8167,
191
- "step": 26
192
- },
193
- {
194
- "epoch": 0.54,
195
- "grad_norm": 0.6043046712875366,
196
- "learning_rate": 0.0002,
197
- "loss": 1.7217,
198
- "step": 27
199
- },
200
- {
201
- "epoch": 0.56,
202
- "grad_norm": 0.6395452618598938,
203
- "learning_rate": 0.0002,
204
- "loss": 1.7353,
205
- "step": 28
206
- },
207
- {
208
- "epoch": 0.58,
209
- "grad_norm": 0.6829009056091309,
210
- "learning_rate": 0.0002,
211
- "loss": 1.7708,
212
- "step": 29
213
- },
214
- {
215
- "epoch": 0.6,
216
- "grad_norm": 0.8561712503433228,
217
- "learning_rate": 0.0002,
218
- "loss": 1.774,
219
- "step": 30
220
- },
221
- {
222
- "epoch": 0.62,
223
- "grad_norm": 0.7594190239906311,
224
- "learning_rate": 0.0002,
225
- "loss": 1.8788,
226
- "step": 31
227
- },
228
- {
229
- "epoch": 0.64,
230
- "grad_norm": 0.867341160774231,
231
- "learning_rate": 0.0002,
232
- "loss": 1.8708,
233
- "step": 32
234
- },
235
- {
236
- "epoch": 0.66,
237
- "grad_norm": 0.9393973350524902,
238
- "learning_rate": 0.0002,
239
- "loss": 1.9839,
240
- "step": 33
241
- },
242
- {
243
- "epoch": 0.68,
244
- "grad_norm": 1.0540133714675903,
245
- "learning_rate": 0.0002,
246
- "loss": 1.7637,
247
- "step": 34
248
- },
249
- {
250
- "epoch": 0.7,
251
- "grad_norm": 1.2020256519317627,
252
- "learning_rate": 0.0002,
253
- "loss": 1.9187,
254
- "step": 35
255
- },
256
- {
257
- "epoch": 0.72,
258
- "grad_norm": 1.7588919401168823,
259
- "learning_rate": 0.0002,
260
- "loss": 1.5851,
261
- "step": 36
262
- },
263
- {
264
- "epoch": 0.74,
265
- "grad_norm": 0.9404975175857544,
266
- "learning_rate": 0.0002,
267
- "loss": 1.7903,
268
- "step": 37
269
- },
270
- {
271
- "epoch": 0.76,
272
- "grad_norm": 0.7744253873825073,
273
- "learning_rate": 0.0002,
274
- "loss": 1.7476,
275
- "step": 38
276
- },
277
- {
278
- "epoch": 0.78,
279
- "grad_norm": 0.7260447144508362,
280
- "learning_rate": 0.0002,
281
- "loss": 1.5547,
282
- "step": 39
283
- },
284
- {
285
- "epoch": 0.8,
286
- "grad_norm": 0.9214150905609131,
287
- "learning_rate": 0.0002,
288
- "loss": 1.6342,
289
- "step": 40
290
- },
291
- {
292
- "epoch": 0.82,
293
- "grad_norm": 0.834932267665863,
294
- "learning_rate": 0.0002,
295
- "loss": 1.6577,
296
- "step": 41
297
- },
298
- {
299
- "epoch": 0.84,
300
- "grad_norm": 0.8663449883460999,
301
- "learning_rate": 0.0002,
302
- "loss": 1.7133,
303
- "step": 42
304
- },
305
- {
306
- "epoch": 0.86,
307
- "grad_norm": 0.9534509181976318,
308
- "learning_rate": 0.0002,
309
- "loss": 1.7999,
310
- "step": 43
311
- },
312
- {
313
- "epoch": 0.88,
314
- "grad_norm": 1.058899164199829,
315
- "learning_rate": 0.0002,
316
- "loss": 1.5964,
317
- "step": 44
318
- },
319
- {
320
- "epoch": 0.9,
321
- "grad_norm": 1.1835004091262817,
322
- "learning_rate": 0.0002,
323
- "loss": 1.7596,
324
- "step": 45
325
- },
326
- {
327
- "epoch": 0.92,
328
- "grad_norm": 1.2041128873825073,
329
- "learning_rate": 0.0002,
330
- "loss": 1.6529,
331
- "step": 46
332
- },
333
- {
334
- "epoch": 0.94,
335
- "grad_norm": 1.5300588607788086,
336
- "learning_rate": 0.0002,
337
- "loss": 1.7852,
338
- "step": 47
339
- },
340
- {
341
- "epoch": 0.96,
342
- "grad_norm": 1.6037429571151733,
343
- "learning_rate": 0.0002,
344
- "loss": 1.6873,
345
- "step": 48
346
- },
347
- {
348
- "epoch": 0.98,
349
- "grad_norm": 0.6437931656837463,
350
- "learning_rate": 0.0002,
351
- "loss": 1.5729,
352
- "step": 49
353
- },
354
- {
355
- "epoch": 1.0,
356
- "grad_norm": 1.1996123790740967,
357
- "learning_rate": 0.0002,
358
- "loss": 1.5609,
359
- "step": 50
360
- },
361
- {
362
- "epoch": 1.0,
363
- "eval_loss": 1.6186010837554932,
364
- "eval_runtime": 565.3333,
365
- "eval_samples_per_second": 0.708,
366
- "eval_steps_per_second": 0.177,
367
- "step": 50
368
- },
369
- {
370
- "epoch": 1.02,
371
- "grad_norm": 0.6480159163475037,
372
- "learning_rate": 0.0002,
373
- "loss": 1.5754,
374
- "step": 51
375
- },
376
- {
377
- "epoch": 1.04,
378
- "grad_norm": 0.8269457817077637,
379
- "learning_rate": 0.0002,
380
- "loss": 1.4743,
381
- "step": 52
382
- },
383
- {
384
- "epoch": 1.06,
385
- "grad_norm": 1.1054280996322632,
386
- "learning_rate": 0.0002,
387
- "loss": 1.3447,
388
- "step": 53
389
- },
390
- {
391
- "epoch": 1.08,
392
- "grad_norm": 0.9144314527511597,
393
- "learning_rate": 0.0002,
394
- "loss": 1.3652,
395
- "step": 54
396
- },
397
- {
398
- "epoch": 1.1,
399
- "grad_norm": 0.8429620862007141,
400
- "learning_rate": 0.0002,
401
- "loss": 1.4696,
402
- "step": 55
403
- },
404
- {
405
- "epoch": 1.12,
406
- "grad_norm": 1.3091776371002197,
407
- "learning_rate": 0.0002,
408
- "loss": 1.3109,
409
- "step": 56
410
- },
411
- {
412
- "epoch": 1.1400000000000001,
413
- "grad_norm": 1.2086460590362549,
414
- "learning_rate": 0.0002,
415
- "loss": 1.3424,
416
- "step": 57
417
- },
418
- {
419
- "epoch": 1.16,
420
- "grad_norm": 1.1823766231536865,
421
- "learning_rate": 0.0002,
422
- "loss": 1.213,
423
- "step": 58
424
- },
425
- {
426
- "epoch": 1.18,
427
- "grad_norm": 1.817803144454956,
428
- "learning_rate": 0.0002,
429
- "loss": 1.2911,
430
- "step": 59
431
- },
432
- {
433
- "epoch": 1.2,
434
- "grad_norm": 1.2870073318481445,
435
- "learning_rate": 0.0002,
436
- "loss": 1.2712,
437
- "step": 60
438
- },
439
- {
440
- "epoch": 1.22,
441
- "grad_norm": 1.2424544095993042,
442
- "learning_rate": 0.0002,
443
- "loss": 1.2292,
444
- "step": 61
445
- },
446
- {
447
- "epoch": 1.24,
448
- "grad_norm": 1.4258471727371216,
449
- "learning_rate": 0.0002,
450
- "loss": 1.0814,
451
- "step": 62
452
- },
453
- {
454
- "epoch": 1.26,
455
- "grad_norm": 1.1297271251678467,
456
- "learning_rate": 0.0002,
457
- "loss": 1.5303,
458
- "step": 63
459
- },
460
- {
461
- "epoch": 1.28,
462
- "grad_norm": 0.8728504776954651,
463
- "learning_rate": 0.0002,
464
- "loss": 1.4354,
465
- "step": 64
466
- },
467
- {
468
- "epoch": 1.3,
469
- "grad_norm": 0.7809789776802063,
470
- "learning_rate": 0.0002,
471
- "loss": 1.4398,
472
- "step": 65
473
- },
474
- {
475
- "epoch": 1.32,
476
- "grad_norm": 0.844166100025177,
477
- "learning_rate": 0.0002,
478
- "loss": 1.2171,
479
- "step": 66
480
- },
481
- {
482
- "epoch": 1.34,
483
- "grad_norm": 0.8636218905448914,
484
- "learning_rate": 0.0002,
485
- "loss": 1.23,
486
- "step": 67
487
- },
488
- {
489
- "epoch": 1.3599999999999999,
490
- "grad_norm": 0.9831591248512268,
491
- "learning_rate": 0.0002,
492
- "loss": 1.2496,
493
- "step": 68
494
- },
495
- {
496
- "epoch": 1.38,
497
- "grad_norm": 1.4268325567245483,
498
- "learning_rate": 0.0002,
499
- "loss": 1.2354,
500
- "step": 69
501
- },
502
- {
503
- "epoch": 1.4,
504
- "grad_norm": 1.6133723258972168,
505
- "learning_rate": 0.0002,
506
- "loss": 1.2123,
507
- "step": 70
508
- },
509
- {
510
- "epoch": 1.42,
511
- "grad_norm": 1.5462720394134521,
512
- "learning_rate": 0.0002,
513
- "loss": 1.0258,
514
- "step": 71
515
- },
516
- {
517
- "epoch": 1.44,
518
- "grad_norm": 1.1962395906448364,
519
- "learning_rate": 0.0002,
520
- "loss": 1.0858,
521
- "step": 72
522
- },
523
- {
524
- "epoch": 1.46,
525
- "grad_norm": 1.413921594619751,
526
- "learning_rate": 0.0002,
527
- "loss": 1.0169,
528
- "step": 73
529
- },
530
- {
531
- "epoch": 1.48,
532
- "grad_norm": 1.442657470703125,
533
- "learning_rate": 0.0002,
534
- "loss": 0.9553,
535
- "step": 74
536
- },
537
- {
538
- "epoch": 1.5,
539
- "grad_norm": 0.9919085502624512,
540
- "learning_rate": 0.0002,
541
- "loss": 1.2394,
542
- "step": 75
543
- },
544
- {
545
- "epoch": 1.52,
546
- "grad_norm": 0.988468587398529,
547
- "learning_rate": 0.0002,
548
- "loss": 1.3642,
549
- "step": 76
550
- },
551
- {
552
- "epoch": 1.54,
553
- "grad_norm": 0.9793186187744141,
554
- "learning_rate": 0.0002,
555
- "loss": 1.2818,
556
- "step": 77
557
- },
558
- {
559
- "epoch": 1.56,
560
- "grad_norm": 0.7799855470657349,
561
- "learning_rate": 0.0002,
562
- "loss": 1.1705,
563
- "step": 78
564
- },
565
- {
566
- "epoch": 1.58,
567
- "grad_norm": 0.8288784027099609,
568
- "learning_rate": 0.0002,
569
- "loss": 1.0484,
570
- "step": 79
571
- },
572
- {
573
- "epoch": 1.6,
574
- "grad_norm": 1.064773440361023,
575
- "learning_rate": 0.0002,
576
- "loss": 1.1608,
577
- "step": 80
578
- },
579
- {
580
- "epoch": 1.62,
581
- "grad_norm": 1.0099600553512573,
582
- "learning_rate": 0.0002,
583
- "loss": 1.1873,
584
- "step": 81
585
- },
586
- {
587
- "epoch": 1.6400000000000001,
588
- "grad_norm": 1.9040124416351318,
589
- "learning_rate": 0.0002,
590
- "loss": 0.9739,
591
- "step": 82
592
- },
593
- {
594
- "epoch": 1.6600000000000001,
595
- "grad_norm": 1.2448644638061523,
596
- "learning_rate": 0.0002,
597
- "loss": 0.9418,
598
- "step": 83
599
- },
600
- {
601
- "epoch": 1.6800000000000002,
602
- "grad_norm": 1.2129086256027222,
603
- "learning_rate": 0.0002,
604
- "loss": 0.8821,
605
- "step": 84
606
- },
607
- {
608
- "epoch": 1.7,
609
- "grad_norm": 1.6727265119552612,
610
- "learning_rate": 0.0002,
611
- "loss": 0.9965,
612
- "step": 85
613
- },
614
- {
615
- "epoch": 1.72,
616
- "grad_norm": 1.6569440364837646,
617
- "learning_rate": 0.0002,
618
- "loss": 0.9182,
619
- "step": 86
620
- },
621
- {
622
- "epoch": 1.74,
623
- "grad_norm": 0.8596146702766418,
624
- "learning_rate": 0.0002,
625
- "loss": 1.3188,
626
- "step": 87
627
- },
628
- {
629
- "epoch": 1.76,
630
- "grad_norm": 0.8928490281105042,
631
- "learning_rate": 0.0002,
632
- "loss": 1.3601,
633
- "step": 88
634
- },
635
- {
636
- "epoch": 1.78,
637
- "grad_norm": 0.7409713268280029,
638
- "learning_rate": 0.0002,
639
- "loss": 1.1212,
640
- "step": 89
641
- },
642
- {
643
- "epoch": 1.8,
644
- "grad_norm": 0.8979334831237793,
645
- "learning_rate": 0.0002,
646
- "loss": 1.4162,
647
- "step": 90
648
- },
649
- {
650
- "epoch": 1.8199999999999998,
651
- "grad_norm": 0.979978621006012,
652
- "learning_rate": 0.0002,
653
- "loss": 1.1969,
654
- "step": 91
655
- },
656
- {
657
- "epoch": 1.8399999999999999,
658
- "grad_norm": 0.9733594059944153,
659
- "learning_rate": 0.0002,
660
- "loss": 1.0468,
661
- "step": 92
662
- },
663
- {
664
- "epoch": 1.8599999999999999,
665
- "grad_norm": 0.9226842522621155,
666
- "learning_rate": 0.0002,
667
- "loss": 1.1807,
668
- "step": 93
669
- },
670
- {
671
- "epoch": 1.88,
672
- "grad_norm": 1.1638745069503784,
673
- "learning_rate": 0.0002,
674
- "loss": 1.139,
675
- "step": 94
676
- },
677
- {
678
- "epoch": 1.9,
679
- "grad_norm": 1.5604937076568604,
680
- "learning_rate": 0.0002,
681
- "loss": 1.1872,
682
- "step": 95
683
- },
684
- {
685
- "epoch": 1.92,
686
- "grad_norm": 1.3674428462982178,
687
- "learning_rate": 0.0002,
688
- "loss": 1.1865,
689
- "step": 96
690
- },
691
- {
692
- "epoch": 1.94,
693
- "grad_norm": 1.8469598293304443,
694
- "learning_rate": 0.0002,
695
- "loss": 1.0469,
696
- "step": 97
697
- },
698
- {
699
- "epoch": 1.96,
700
- "grad_norm": 1.3148952722549438,
701
- "learning_rate": 0.0002,
702
- "loss": 0.9915,
703
- "step": 98
704
- },
705
- {
706
- "epoch": 1.98,
707
- "grad_norm": 1.599141001701355,
708
- "learning_rate": 0.0002,
709
- "loss": 1.2296,
710
- "step": 99
711
- },
712
- {
713
- "epoch": 2.0,
714
- "grad_norm": 1.3382114171981812,
715
- "learning_rate": 0.0002,
716
- "loss": 1.1813,
717
- "step": 100
718
- },
719
- {
720
- "epoch": 2.0,
721
- "eval_loss": 1.409305453300476,
722
- "eval_runtime": 565.9517,
723
- "eval_samples_per_second": 0.707,
724
- "eval_steps_per_second": 0.177,
725
- "step": 100
726
- },
727
- {
728
- "epoch": 2.02,
729
- "grad_norm": 1.0162380933761597,
730
- "learning_rate": 0.0002,
731
- "loss": 1.1481,
732
- "step": 101
733
- },
734
- {
735
- "epoch": 2.04,
736
- "grad_norm": 0.7402092814445496,
737
- "learning_rate": 0.0002,
738
- "loss": 1.0086,
739
- "step": 102
740
- },
741
- {
742
- "epoch": 2.06,
743
- "grad_norm": 0.8824872970581055,
744
- "learning_rate": 0.0002,
745
- "loss": 1.0588,
746
- "step": 103
747
- },
748
- {
749
- "epoch": 2.08,
750
- "grad_norm": 0.7582442760467529,
751
- "learning_rate": 0.0002,
752
- "loss": 0.8181,
753
- "step": 104
754
- },
755
- {
756
- "epoch": 2.1,
757
- "grad_norm": 1.0200812816619873,
758
- "learning_rate": 0.0002,
759
- "loss": 0.9041,
760
- "step": 105
761
- },
762
- {
763
- "epoch": 2.12,
764
- "grad_norm": 1.08174467086792,
765
- "learning_rate": 0.0002,
766
- "loss": 0.8479,
767
- "step": 106
768
- },
769
- {
770
- "epoch": 2.14,
771
- "grad_norm": 1.01225745677948,
772
- "learning_rate": 0.0002,
773
- "loss": 0.7659,
774
- "step": 107
775
- },
776
- {
777
- "epoch": 2.16,
778
- "grad_norm": 1.2194840908050537,
779
- "learning_rate": 0.0002,
780
- "loss": 0.7926,
781
- "step": 108
782
- },
783
- {
784
- "epoch": 2.18,
785
- "grad_norm": 1.0519524812698364,
786
- "learning_rate": 0.0002,
787
- "loss": 0.6604,
788
- "step": 109
789
- },
790
- {
791
- "epoch": 2.2,
792
- "grad_norm": 1.2860150337219238,
793
- "learning_rate": 0.0002,
794
- "loss": 0.663,
795
- "step": 110
796
- },
797
- {
798
- "epoch": 2.22,
799
- "grad_norm": 1.5521994829177856,
800
- "learning_rate": 0.0002,
801
- "loss": 0.7791,
802
- "step": 111
803
- },
804
- {
805
- "epoch": 2.24,
806
- "grad_norm": 1.455283284187317,
807
- "learning_rate": 0.0002,
808
- "loss": 0.5112,
809
- "step": 112
810
- },
811
- {
812
- "epoch": 2.26,
813
- "grad_norm": 1.7097278833389282,
814
- "learning_rate": 0.0002,
815
- "loss": 1.2219,
816
- "step": 113
817
- },
818
- {
819
- "epoch": 2.2800000000000002,
820
- "grad_norm": 1.5385531187057495,
821
- "learning_rate": 0.0002,
822
- "loss": 1.1261,
823
- "step": 114
824
- },
825
- {
826
- "epoch": 2.3,
827
- "grad_norm": 1.0525436401367188,
828
- "learning_rate": 0.0002,
829
- "loss": 0.86,
830
- "step": 115
831
- },
832
- {
833
- "epoch": 2.32,
834
- "grad_norm": 1.0388120412826538,
835
- "learning_rate": 0.0002,
836
- "loss": 0.9022,
837
- "step": 116
838
- },
839
- {
840
- "epoch": 2.34,
841
- "grad_norm": 1.060497760772705,
842
- "learning_rate": 0.0002,
843
- "loss": 0.9265,
844
- "step": 117
845
- },
846
- {
847
- "epoch": 2.36,
848
- "grad_norm": 1.0629950761795044,
849
- "learning_rate": 0.0002,
850
- "loss": 0.7222,
851
- "step": 118
852
- },
853
- {
854
- "epoch": 2.38,
855
- "grad_norm": 1.2574018239974976,
856
- "learning_rate": 0.0002,
857
- "loss": 0.7952,
858
- "step": 119
859
- },
860
- {
861
- "epoch": 2.4,
862
- "grad_norm": 1.0951610803604126,
863
- "learning_rate": 0.0002,
864
- "loss": 0.6647,
865
- "step": 120
866
- },
867
- {
868
- "epoch": 2.42,
869
- "grad_norm": 1.46285879611969,
870
- "learning_rate": 0.0002,
871
- "loss": 0.7845,
872
- "step": 121
873
- },
874
- {
875
- "epoch": 2.44,
876
- "grad_norm": 1.3611388206481934,
877
- "learning_rate": 0.0002,
878
- "loss": 0.7084,
879
- "step": 122
880
- },
881
- {
882
- "epoch": 2.46,
883
- "grad_norm": 1.6670907735824585,
884
- "learning_rate": 0.0002,
885
- "loss": 0.6594,
886
- "step": 123
887
- },
888
- {
889
- "epoch": 2.48,
890
- "grad_norm": 2.1525955200195312,
891
- "learning_rate": 0.0002,
892
- "loss": 0.6401,
893
- "step": 124
894
- },
895
- {
896
- "epoch": 2.5,
897
- "grad_norm": 2.5126793384552,
898
- "learning_rate": 0.0002,
899
- "loss": 0.926,
900
- "step": 125
901
- },
902
- {
903
- "epoch": 2.52,
904
- "grad_norm": 1.800521969795227,
905
- "learning_rate": 0.0002,
906
- "loss": 1.0936,
907
- "step": 126
908
- },
909
- {
910
- "epoch": 2.54,
911
- "grad_norm": 1.0617576837539673,
912
- "learning_rate": 0.0002,
913
- "loss": 0.8052,
914
- "step": 127
915
- },
916
- {
917
- "epoch": 2.56,
918
- "grad_norm": 1.0823312997817993,
919
- "learning_rate": 0.0002,
920
- "loss": 0.9443,
921
- "step": 128
922
- },
923
- {
924
- "epoch": 2.58,
925
- "grad_norm": 1.2193264961242676,
926
- "learning_rate": 0.0002,
927
- "loss": 0.7955,
928
- "step": 129
929
- },
930
- {
931
- "epoch": 2.6,
932
- "grad_norm": 1.0502954721450806,
933
- "learning_rate": 0.0002,
934
- "loss": 0.8365,
935
- "step": 130
936
- },
937
- {
938
- "epoch": 2.62,
939
- "grad_norm": 1.1898560523986816,
940
- "learning_rate": 0.0002,
941
- "loss": 0.9706,
942
- "step": 131
943
- },
944
- {
945
- "epoch": 2.64,
946
- "grad_norm": 1.1076680421829224,
947
- "learning_rate": 0.0002,
948
- "loss": 0.7529,
949
- "step": 132
950
- },
951
- {
952
- "epoch": 2.66,
953
- "grad_norm": 1.3826709985733032,
954
- "learning_rate": 0.0002,
955
- "loss": 0.7474,
956
- "step": 133
957
- },
958
- {
959
- "epoch": 2.68,
960
- "grad_norm": 1.2504832744598389,
961
- "learning_rate": 0.0002,
962
- "loss": 0.7086,
963
- "step": 134
964
- },
965
- {
966
- "epoch": 2.7,
967
- "grad_norm": 1.6292765140533447,
968
- "learning_rate": 0.0002,
969
- "loss": 0.6305,
970
- "step": 135
971
- },
972
- {
973
- "epoch": 2.7199999999999998,
974
- "grad_norm": 1.9603074789047241,
975
- "learning_rate": 0.0002,
976
- "loss": 0.6834,
977
- "step": 136
978
- },
979
- {
980
- "epoch": 2.74,
981
- "grad_norm": 2.202030897140503,
982
- "learning_rate": 0.0002,
983
- "loss": 1.1712,
984
- "step": 137
985
- },
986
- {
987
- "epoch": 2.76,
988
- "grad_norm": 1.6344685554504395,
989
- "learning_rate": 0.0002,
990
- "loss": 1.0772,
991
- "step": 138
992
- },
993
- {
994
- "epoch": 2.7800000000000002,
995
- "grad_norm": 1.3579537868499756,
996
- "learning_rate": 0.0002,
997
- "loss": 0.8803,
998
- "step": 139
999
- },
1000
- {
1001
- "epoch": 2.8,
1002
- "grad_norm": 1.0554553270339966,
1003
- "learning_rate": 0.0002,
1004
- "loss": 0.9222,
1005
- "step": 140
1006
- },
1007
- {
1008
- "epoch": 2.82,
1009
- "grad_norm": 0.9431642889976501,
1010
- "learning_rate": 0.0002,
1011
- "loss": 0.8031,
1012
- "step": 141
1013
- },
1014
- {
1015
- "epoch": 2.84,
1016
- "grad_norm": 1.0826098918914795,
1017
- "learning_rate": 0.0002,
1018
- "loss": 0.8259,
1019
- "step": 142
1020
- },
1021
- {
1022
- "epoch": 2.86,
1023
- "grad_norm": 1.24959135055542,
1024
- "learning_rate": 0.0002,
1025
- "loss": 0.7957,
1026
- "step": 143
1027
- },
1028
- {
1029
- "epoch": 2.88,
1030
- "grad_norm": 1.1057368516921997,
1031
- "learning_rate": 0.0002,
1032
- "loss": 0.7079,
1033
- "step": 144
1034
- },
1035
- {
1036
- "epoch": 2.9,
1037
- "grad_norm": 1.144061803817749,
1038
- "learning_rate": 0.0002,
1039
- "loss": 0.7165,
1040
- "step": 145
1041
- },
1042
- {
1043
- "epoch": 2.92,
1044
- "grad_norm": 1.0690631866455078,
1045
- "learning_rate": 0.0002,
1046
- "loss": 0.601,
1047
- "step": 146
1048
- },
1049
- {
1050
- "epoch": 2.94,
1051
- "grad_norm": 1.292758584022522,
1052
- "learning_rate": 0.0002,
1053
- "loss": 0.7191,
1054
- "step": 147
1055
- },
1056
- {
1057
- "epoch": 2.96,
1058
- "grad_norm": 1.729408860206604,
1059
- "learning_rate": 0.0002,
1060
- "loss": 0.5851,
1061
- "step": 148
1062
- },
1063
- {
1064
- "epoch": 2.98,
1065
- "grad_norm": 2.078197717666626,
1066
- "learning_rate": 0.0002,
1067
- "loss": 0.8942,
1068
- "step": 149
1069
- },
1070
- {
1071
- "epoch": 3.0,
1072
- "grad_norm": 2.0007128715515137,
1073
- "learning_rate": 0.0002,
1074
- "loss": 0.7139,
1075
- "step": 150
1076
- },
1077
- {
1078
- "epoch": 3.0,
1079
- "eval_loss": 1.4874461889266968,
1080
- "eval_runtime": 566.1506,
1081
- "eval_samples_per_second": 0.707,
1082
- "eval_steps_per_second": 0.177,
1083
- "step": 150
1084
- },
1085
- {
1086
- "epoch": 3.02,
1087
- "grad_norm": 1.024670958518982,
1088
- "learning_rate": 0.0002,
1089
- "loss": 0.8969,
1090
- "step": 151
1091
- },
1092
- {
1093
- "epoch": 3.04,
1094
- "grad_norm": 0.9738882184028625,
1095
- "learning_rate": 0.0002,
1096
- "loss": 0.9056,
1097
- "step": 152
1098
- },
1099
- {
1100
- "epoch": 3.06,
1101
- "grad_norm": 0.9969688653945923,
1102
- "learning_rate": 0.0002,
1103
- "loss": 0.6676,
1104
- "step": 153
1105
- },
1106
- {
1107
- "epoch": 3.08,
1108
- "grad_norm": 1.12136971950531,
1109
- "learning_rate": 0.0002,
1110
- "loss": 0.5916,
1111
- "step": 154
1112
- },
1113
- {
1114
- "epoch": 3.1,
1115
- "grad_norm": 1.3517699241638184,
1116
- "learning_rate": 0.0002,
1117
- "loss": 0.5908,
1118
- "step": 155
1119
- },
1120
- {
1121
- "epoch": 3.12,
1122
- "grad_norm": 1.5965360403060913,
1123
- "learning_rate": 0.0002,
1124
- "loss": 0.6148,
1125
- "step": 156
1126
- },
1127
- {
1128
- "epoch": 3.14,
1129
- "grad_norm": 1.3009252548217773,
1130
- "learning_rate": 0.0002,
1131
- "loss": 0.4902,
1132
- "step": 157
1133
- },
1134
- {
1135
- "epoch": 3.16,
1136
- "grad_norm": 1.2742400169372559,
1137
- "learning_rate": 0.0002,
1138
- "loss": 0.4515,
1139
- "step": 158
1140
- },
1141
- {
1142
- "epoch": 3.18,
1143
- "grad_norm": 1.2994771003723145,
1144
- "learning_rate": 0.0002,
1145
- "loss": 0.416,
1146
- "step": 159
1147
- },
1148
- {
1149
- "epoch": 3.2,
1150
- "grad_norm": 1.3306324481964111,
1151
- "learning_rate": 0.0002,
1152
- "loss": 0.4144,
1153
- "step": 160
1154
- },
1155
- {
1156
- "epoch": 3.22,
1157
- "grad_norm": 1.5406475067138672,
1158
- "learning_rate": 0.0002,
1159
- "loss": 0.4256,
1160
- "step": 161
1161
- },
1162
- {
1163
- "epoch": 3.24,
1164
- "grad_norm": 1.584506630897522,
1165
- "learning_rate": 0.0002,
1166
- "loss": 0.4511,
1167
- "step": 162
1168
- },
1169
- {
1170
- "epoch": 3.26,
1171
- "grad_norm": 1.6618622541427612,
1172
- "learning_rate": 0.0002,
1173
- "loss": 0.8865,
1174
- "step": 163
1175
- },
1176
- {
1177
- "epoch": 3.2800000000000002,
1178
- "grad_norm": 1.6019847393035889,
1179
- "learning_rate": 0.0002,
1180
- "loss": 0.7339,
1181
- "step": 164
1182
- },
1183
- {
1184
- "epoch": 3.3,
1185
- "grad_norm": 1.1740251779556274,
1186
- "learning_rate": 0.0002,
1187
- "loss": 0.6945,
1188
- "step": 165
1189
- },
1190
- {
1191
- "epoch": 3.32,
1192
- "grad_norm": 1.1268410682678223,
1193
- "learning_rate": 0.0002,
1194
- "loss": 0.579,
1195
- "step": 166
1196
- },
1197
- {
1198
- "epoch": 3.34,
1199
- "grad_norm": 1.3038002252578735,
1200
- "learning_rate": 0.0002,
1201
- "loss": 0.5217,
1202
- "step": 167
1203
- },
1204
- {
1205
- "epoch": 3.36,
1206
- "grad_norm": 1.112185001373291,
1207
- "learning_rate": 0.0002,
1208
- "loss": 0.4766,
1209
- "step": 168
1210
- },
1211
- {
1212
- "epoch": 3.38,
1213
- "grad_norm": 1.3828542232513428,
1214
- "learning_rate": 0.0002,
1215
- "loss": 0.4781,
1216
- "step": 169
1217
- },
1218
- {
1219
- "epoch": 3.4,
1220
- "grad_norm": 1.1456600427627563,
1221
- "learning_rate": 0.0002,
1222
- "loss": 0.4056,
1223
- "step": 170
1224
- },
1225
- {
1226
- "epoch": 3.42,
1227
- "grad_norm": 1.2479093074798584,
1228
- "learning_rate": 0.0002,
1229
- "loss": 0.4447,
1230
- "step": 171
1231
- },
1232
- {
1233
- "epoch": 3.44,
1234
- "grad_norm": 1.4044010639190674,
1235
- "learning_rate": 0.0002,
1236
- "loss": 0.3814,
1237
- "step": 172
1238
- },
1239
- {
1240
- "epoch": 3.46,
1241
- "grad_norm": 1.565138339996338,
1242
- "learning_rate": 0.0002,
1243
- "loss": 0.3982,
1244
- "step": 173
1245
- },
1246
- {
1247
- "epoch": 3.48,
1248
- "grad_norm": 1.4442418813705444,
1249
- "learning_rate": 0.0002,
1250
- "loss": 0.4262,
1251
- "step": 174
1252
- },
1253
- {
1254
- "epoch": 3.5,
1255
- "grad_norm": 1.1203701496124268,
1256
- "learning_rate": 0.0002,
1257
- "loss": 0.8025,
1258
- "step": 175
1259
- },
1260
- {
1261
- "epoch": 3.52,
1262
- "grad_norm": 1.3620504140853882,
1263
- "learning_rate": 0.0002,
1264
- "loss": 0.9045,
1265
- "step": 176
1266
- },
1267
- {
1268
- "epoch": 3.54,
1269
- "grad_norm": 1.5145343542099,
1270
- "learning_rate": 0.0002,
1271
- "loss": 0.703,
1272
- "step": 177
1273
- },
1274
- {
1275
- "epoch": 3.56,
1276
- "grad_norm": 1.333682656288147,
1277
- "learning_rate": 0.0002,
1278
- "loss": 0.6285,
1279
- "step": 178
1280
- },
1281
- {
1282
- "epoch": 3.58,
1283
- "grad_norm": 1.4228661060333252,
1284
- "learning_rate": 0.0002,
1285
- "loss": 0.6004,
1286
- "step": 179
1287
- },
1288
- {
1289
- "epoch": 3.6,
1290
- "grad_norm": 1.2111386060714722,
1291
- "learning_rate": 0.0002,
1292
- "loss": 0.4754,
1293
- "step": 180
1294
- },
1295
- {
1296
- "epoch": 3.62,
1297
- "grad_norm": 1.410719394683838,
1298
- "learning_rate": 0.0002,
1299
- "loss": 0.5324,
1300
- "step": 181
1301
- },
1302
- {
1303
- "epoch": 3.64,
1304
- "grad_norm": 1.4157259464263916,
1305
- "learning_rate": 0.0002,
1306
- "loss": 0.4556,
1307
- "step": 182
1308
- },
1309
- {
1310
- "epoch": 3.66,
1311
- "grad_norm": 1.3982216119766235,
1312
- "learning_rate": 0.0002,
1313
- "loss": 0.4465,
1314
- "step": 183
1315
- },
1316
- {
1317
- "epoch": 3.68,
1318
- "grad_norm": 1.4364334344863892,
1319
- "learning_rate": 0.0002,
1320
- "loss": 0.4313,
1321
- "step": 184
1322
- },
1323
- {
1324
- "epoch": 3.7,
1325
- "grad_norm": 1.5408861637115479,
1326
- "learning_rate": 0.0002,
1327
- "loss": 0.4108,
1328
- "step": 185
1329
- },
1330
- {
1331
- "epoch": 3.7199999999999998,
1332
- "grad_norm": 1.5500551462173462,
1333
- "learning_rate": 0.0002,
1334
- "loss": 0.4664,
1335
- "step": 186
1336
- },
1337
- {
1338
- "epoch": 3.74,
1339
- "grad_norm": 1.1150060892105103,
1340
- "learning_rate": 0.0002,
1341
- "loss": 0.8951,
1342
- "step": 187
1343
- },
1344
- {
1345
- "epoch": 3.76,
1346
- "grad_norm": 1.0168464183807373,
1347
- "learning_rate": 0.0002,
1348
- "loss": 0.8619,
1349
- "step": 188
1350
- },
1351
- {
1352
- "epoch": 3.7800000000000002,
1353
- "grad_norm": 1.2093026638031006,
1354
- "learning_rate": 0.0002,
1355
- "loss": 0.5782,
1356
- "step": 189
1357
- },
1358
- {
1359
- "epoch": 3.8,
1360
- "grad_norm": 1.3905984163284302,
1361
- "learning_rate": 0.0002,
1362
- "loss": 0.6929,
1363
- "step": 190
1364
- },
1365
- {
1366
- "epoch": 3.82,
1367
- "grad_norm": 1.3665902614593506,
1368
- "learning_rate": 0.0002,
1369
- "loss": 0.5701,
1370
- "step": 191
1371
- },
1372
- {
1373
- "epoch": 3.84,
1374
- "grad_norm": 1.1478445529937744,
1375
- "learning_rate": 0.0002,
1376
- "loss": 0.4515,
1377
- "step": 192
1378
- },
1379
- {
1380
- "epoch": 3.86,
1381
- "grad_norm": 1.2758458852767944,
1382
- "learning_rate": 0.0002,
1383
- "loss": 0.5819,
1384
- "step": 193
1385
- },
1386
- {
1387
- "epoch": 3.88,
1388
- "grad_norm": 1.0731329917907715,
1389
- "learning_rate": 0.0002,
1390
- "loss": 0.4448,
1391
- "step": 194
1392
- },
1393
- {
1394
- "epoch": 3.9,
1395
- "grad_norm": 1.20659339427948,
1396
- "learning_rate": 0.0002,
1397
- "loss": 0.4295,
1398
- "step": 195
1399
- },
1400
- {
1401
- "epoch": 3.92,
1402
- "grad_norm": 1.3976835012435913,
1403
- "learning_rate": 0.0002,
1404
- "loss": 0.4566,
1405
- "step": 196
1406
- },
1407
- {
1408
- "epoch": 3.94,
1409
- "grad_norm": 1.617711067199707,
1410
- "learning_rate": 0.0002,
1411
- "loss": 0.4303,
1412
- "step": 197
1413
- },
1414
- {
1415
- "epoch": 3.96,
1416
- "grad_norm": 1.707471489906311,
1417
- "learning_rate": 0.0002,
1418
- "loss": 0.4539,
1419
- "step": 198
1420
- },
1421
- {
1422
- "epoch": 3.98,
1423
- "grad_norm": 1.2962028980255127,
1424
- "learning_rate": 0.0002,
1425
- "loss": 0.5971,
1426
- "step": 199
1427
- },
1428
- {
1429
- "epoch": 4.0,
1430
- "grad_norm": 1.8809109926223755,
1431
- "learning_rate": 0.0002,
1432
- "loss": 0.4973,
1433
- "step": 200
1434
- },
1435
- {
1436
- "epoch": 4.0,
1437
- "eval_loss": 1.5414550304412842,
1438
- "eval_runtime": 565.7272,
1439
- "eval_samples_per_second": 0.707,
1440
- "eval_steps_per_second": 0.177,
1441
- "step": 200
1442
- },
1443
- {
1444
- "epoch": 4.02,
1445
- "grad_norm": 0.9540490508079529,
1446
- "learning_rate": 0.0002,
1447
- "loss": 0.56,
1448
- "step": 201
1449
- },
1450
- {
1451
- "epoch": 4.04,
1452
- "grad_norm": 1.0443426370620728,
1453
- "learning_rate": 0.0002,
1454
- "loss": 0.6562,
1455
- "step": 202
1456
- },
1457
- {
1458
- "epoch": 4.06,
1459
- "grad_norm": 1.020203948020935,
1460
- "learning_rate": 0.0002,
1461
- "loss": 0.4533,
1462
- "step": 203
1463
- },
1464
- {
1465
- "epoch": 4.08,
1466
- "grad_norm": 1.5309128761291504,
1467
- "learning_rate": 0.0002,
1468
- "loss": 0.3575,
1469
- "step": 204
1470
- },
1471
- {
1472
- "epoch": 4.1,
1473
- "grad_norm": 1.7135676145553589,
1474
- "learning_rate": 0.0002,
1475
- "loss": 0.3286,
1476
- "step": 205
1477
- },
1478
- {
1479
- "epoch": 4.12,
1480
- "grad_norm": 1.602728247642517,
1481
- "learning_rate": 0.0002,
1482
- "loss": 0.2556,
1483
- "step": 206
1484
- },
1485
- {
1486
- "epoch": 4.14,
1487
- "grad_norm": 1.8623350858688354,
1488
- "learning_rate": 0.0002,
1489
- "loss": 0.314,
1490
- "step": 207
1491
- },
1492
- {
1493
- "epoch": 4.16,
1494
- "grad_norm": 1.5630223751068115,
1495
- "learning_rate": 0.0002,
1496
- "loss": 0.2716,
1497
- "step": 208
1498
- },
1499
- {
1500
- "epoch": 4.18,
1501
- "grad_norm": 1.3671077489852905,
1502
- "learning_rate": 0.0002,
1503
- "loss": 0.2506,
1504
- "step": 209
1505
- },
1506
- {
1507
- "epoch": 4.2,
1508
- "grad_norm": 1.0884723663330078,
1509
- "learning_rate": 0.0002,
1510
- "loss": 0.2473,
1511
- "step": 210
1512
- },
1513
- {
1514
- "epoch": 4.22,
1515
- "grad_norm": 1.193832516670227,
1516
- "learning_rate": 0.0002,
1517
- "loss": 0.2836,
1518
- "step": 211
1519
- },
1520
- {
1521
- "epoch": 4.24,
1522
- "grad_norm": 1.0041422843933105,
1523
- "learning_rate": 0.0002,
1524
- "loss": 0.391,
1525
- "step": 212
1526
- },
1527
- {
1528
- "epoch": 4.26,
1529
- "grad_norm": 1.013597846031189,
1530
- "learning_rate": 0.0002,
1531
- "loss": 0.628,
1532
- "step": 213
1533
- },
1534
- {
1535
- "epoch": 4.28,
1536
- "grad_norm": 0.9650751948356628,
1537
- "learning_rate": 0.0002,
1538
- "loss": 0.5202,
1539
- "step": 214
1540
- },
1541
- {
1542
- "epoch": 4.3,
1543
- "grad_norm": 1.0781069993972778,
1544
- "learning_rate": 0.0002,
1545
- "loss": 0.4967,
1546
- "step": 215
1547
- },
1548
- {
1549
- "epoch": 4.32,
1550
- "grad_norm": 1.1297317743301392,
1551
- "learning_rate": 0.0002,
1552
- "loss": 0.4154,
1553
- "step": 216
1554
- },
1555
- {
1556
- "epoch": 4.34,
1557
- "grad_norm": 1.2913479804992676,
1558
- "learning_rate": 0.0002,
1559
- "loss": 0.3014,
1560
- "step": 217
1561
- },
1562
- {
1563
- "epoch": 4.36,
1564
- "grad_norm": 1.4399878978729248,
1565
- "learning_rate": 0.0002,
1566
- "loss": 0.3344,
1567
- "step": 218
1568
- },
1569
- {
1570
- "epoch": 4.38,
1571
- "grad_norm": 1.4960243701934814,
1572
- "learning_rate": 0.0002,
1573
- "loss": 0.2894,
1574
- "step": 219
1575
- },
1576
- {
1577
- "epoch": 4.4,
1578
- "grad_norm": 1.925826072692871,
1579
- "learning_rate": 0.0002,
1580
- "loss": 0.281,
1581
- "step": 220
1582
- },
1583
- {
1584
- "epoch": 4.42,
1585
- "grad_norm": 1.6930102109909058,
1586
- "learning_rate": 0.0002,
1587
- "loss": 0.2512,
1588
- "step": 221
1589
- },
1590
- {
1591
- "epoch": 4.44,
1592
- "grad_norm": 1.6776522397994995,
1593
- "learning_rate": 0.0002,
1594
- "loss": 0.2744,
1595
- "step": 222
1596
- },
1597
- {
1598
- "epoch": 4.46,
1599
- "grad_norm": 1.3323974609375,
1600
- "learning_rate": 0.0002,
1601
- "loss": 0.2951,
1602
- "step": 223
1603
- },
1604
- {
1605
- "epoch": 4.48,
1606
- "grad_norm": 1.2120009660720825,
1607
- "learning_rate": 0.0002,
1608
- "loss": 0.3707,
1609
- "step": 224
1610
- },
1611
- {
1612
- "epoch": 4.5,
1613
- "grad_norm": 1.2817238569259644,
1614
- "learning_rate": 0.0002,
1615
- "loss": 0.6035,
1616
- "step": 225
1617
- },
1618
- {
1619
- "epoch": 4.52,
1620
- "grad_norm": 1.1797271966934204,
1621
- "learning_rate": 0.0002,
1622
- "loss": 0.5878,
1623
- "step": 226
1624
- },
1625
- {
1626
- "epoch": 4.54,
1627
- "grad_norm": 0.9533390402793884,
1628
- "learning_rate": 0.0002,
1629
- "loss": 0.3458,
1630
- "step": 227
1631
- },
1632
- {
1633
- "epoch": 4.5600000000000005,
1634
- "grad_norm": 1.0915648937225342,
1635
- "learning_rate": 0.0002,
1636
- "loss": 0.3617,
1637
- "step": 228
1638
- },
1639
- {
1640
- "epoch": 4.58,
1641
- "grad_norm": 1.3463889360427856,
1642
- "learning_rate": 0.0002,
1643
- "loss": 0.3853,
1644
- "step": 229
1645
- },
1646
- {
1647
- "epoch": 4.6,
1648
- "grad_norm": 1.457556128501892,
1649
- "learning_rate": 0.0002,
1650
- "loss": 0.3343,
1651
- "step": 230
1652
- },
1653
- {
1654
- "epoch": 4.62,
1655
- "grad_norm": 1.681526780128479,
1656
- "learning_rate": 0.0002,
1657
- "loss": 0.3132,
1658
- "step": 231
1659
- },
1660
- {
1661
- "epoch": 4.64,
1662
- "grad_norm": 1.7101032733917236,
1663
- "learning_rate": 0.0002,
1664
- "loss": 0.2913,
1665
- "step": 232
1666
- },
1667
- {
1668
- "epoch": 4.66,
1669
- "grad_norm": 2.1125667095184326,
1670
- "learning_rate": 0.0002,
1671
- "loss": 0.304,
1672
- "step": 233
1673
- },
1674
- {
1675
- "epoch": 4.68,
1676
- "grad_norm": 1.5824134349822998,
1677
- "learning_rate": 0.0002,
1678
- "loss": 0.2878,
1679
- "step": 234
1680
- },
1681
- {
1682
- "epoch": 4.7,
1683
- "grad_norm": 1.5257948637008667,
1684
- "learning_rate": 0.0002,
1685
- "loss": 0.3049,
1686
- "step": 235
1687
- },
1688
- {
1689
- "epoch": 4.72,
1690
- "grad_norm": 1.1413626670837402,
1691
- "learning_rate": 0.0002,
1692
- "loss": 0.3875,
1693
- "step": 236
1694
- },
1695
- {
1696
- "epoch": 4.74,
1697
- "grad_norm": 1.4785950183868408,
1698
- "learning_rate": 0.0002,
1699
- "loss": 0.7915,
1700
- "step": 237
1701
- },
1702
- {
1703
- "epoch": 4.76,
1704
- "grad_norm": 1.0445829629898071,
1705
- "learning_rate": 0.0002,
1706
- "loss": 0.4765,
1707
- "step": 238
1708
- },
1709
- {
1710
- "epoch": 4.78,
1711
- "grad_norm": 1.0932363271713257,
1712
- "learning_rate": 0.0002,
1713
- "loss": 0.4081,
1714
- "step": 239
1715
- },
1716
- {
1717
- "epoch": 4.8,
1718
- "grad_norm": 1.313068151473999,
1719
- "learning_rate": 0.0002,
1720
- "loss": 0.4278,
1721
- "step": 240
1722
- },
1723
- {
1724
- "epoch": 4.82,
1725
- "grad_norm": 1.2771199941635132,
1726
- "learning_rate": 0.0002,
1727
- "loss": 0.3481,
1728
- "step": 241
1729
- },
1730
- {
1731
- "epoch": 4.84,
1732
- "grad_norm": 1.3306118249893188,
1733
- "learning_rate": 0.0002,
1734
- "loss": 0.3301,
1735
- "step": 242
1736
- },
1737
- {
1738
- "epoch": 4.86,
1739
- "grad_norm": 1.2204334735870361,
1740
- "learning_rate": 0.0002,
1741
- "loss": 0.2974,
1742
- "step": 243
1743
- },
1744
- {
1745
- "epoch": 4.88,
1746
- "grad_norm": 1.1585593223571777,
1747
- "learning_rate": 0.0002,
1748
- "loss": 0.2404,
1749
- "step": 244
1750
- },
1751
- {
1752
- "epoch": 4.9,
1753
- "grad_norm": 1.6888794898986816,
1754
- "learning_rate": 0.0002,
1755
- "loss": 0.2817,
1756
- "step": 245
1757
- },
1758
- {
1759
- "epoch": 4.92,
1760
- "grad_norm": 1.4956034421920776,
1761
- "learning_rate": 0.0002,
1762
- "loss": 0.2583,
1763
- "step": 246
1764
- },
1765
- {
1766
- "epoch": 4.9399999999999995,
1767
- "grad_norm": 1.6130638122558594,
1768
- "learning_rate": 0.0002,
1769
- "loss": 0.2888,
1770
- "step": 247
1771
- },
1772
- {
1773
- "epoch": 4.96,
1774
- "grad_norm": 2.280722141265869,
1775
- "learning_rate": 0.0002,
1776
- "loss": 0.3732,
1777
- "step": 248
1778
- },
1779
- {
1780
- "epoch": 4.98,
1781
- "grad_norm": 1.8177040815353394,
1782
- "learning_rate": 0.0002,
1783
- "loss": 0.6086,
1784
- "step": 249
1785
- },
1786
- {
1787
- "epoch": 5.0,
1788
- "grad_norm": 1.674232840538025,
1789
- "learning_rate": 0.0002,
1790
- "loss": 0.3232,
1791
- "step": 250
1792
- },
1793
- {
1794
- "epoch": 5.0,
1795
- "eval_loss": 1.6614739894866943,
1796
- "eval_runtime": 565.5774,
1797
- "eval_samples_per_second": 0.707,
1798
- "eval_steps_per_second": 0.177,
1799
- "step": 250
1800
- }
1801
- ],
1802
- "logging_steps": 1,
1803
- "max_steps": 300,
1804
- "num_input_tokens_seen": 0,
1805
- "num_train_epochs": 6,
1806
- "save_steps": 500,
1807
- "stateful_callbacks": {
1808
- "TrainerControl": {
1809
- "args": {
1810
- "should_epoch_stop": false,
1811
- "should_evaluate": false,
1812
- "should_log": false,
1813
- "should_save": true,
1814
- "should_training_stop": false
1815
- },
1816
- "attributes": {}
1817
- }
1818
- },
1819
- "total_flos": 9.433614153351168e+16,
1820
- "train_batch_size": 8,
1821
- "trial_name": null,
1822
- "trial_params": null
1823
- }
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
checkpoint-250/training_args.bin DELETED
@@ -1,3 +0,0 @@
1
- version https://git-lfs.github.com/spec/v1
2
- oid sha256:065527ed19cb54e2fb4d35cca4577bb1742be7124caa29ec286008d30a425c18
3
- size 5368