boundary_pretraining¶
topic_segmentation.boundary_pretraining
Pre-train ModernBERT on C4 paragraph boundaries.
Label 1 marks the last token of a paragraph. Settings come from configs/boundary_pretraining.json. Launch with scripts/run_boundary_pretraining.py; training resumes from the last checkpoint in --output-dir.
SaveEachEpoch
¶
Bases: TrainerCallback
Save the weights at the end of every epoch, independent of save_total_limit.
Source code in src/topic_segmentation/boundary_pretraining.py
25 26 27 28 29 30 31 32 33 34 35 | |
tokenizer = tokenizer
instance-attribute
¶
on_epoch_end(args, state, control, model=None, **kwargs)
¶
Source code in src/topic_segmentation/boundary_pretraining.py
31 32 33 34 35 | |
boundary_scores(prediction)
¶
Token-level precision, recall, and F1 of the paragraph-end class.
Source code in src/topic_segmentation/boundary_pretraining.py
38 39 40 41 42 43 | |
load_model(base, seed, attention)
¶
Load base with a new token-classification head; seeding first fixes the head's initialization.
Source code in src/topic_segmentation/boundary_pretraining.py
46 47 48 49 50 51 | |
training_arguments(training, output_dir, run_name)
¶
Build Trainer arguments from the training configuration.
Source code in src/topic_segmentation/boundary_pretraining.py
54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 | |
main()
¶
Source code in src/topic_segmentation/boundary_pretraining.py
85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | |