Topic Segmentation¶
Code for news topic segmentation and transfer to clinical notes.
Read the online documentation for the user guide and API reference.
See SETUP.md for installation, downloads, and parameter examples.
The models mark section boundaries between paragraphs or sentences.
The final model uses ModernBERT-base with C4 boundary pre-training, MIND adaptation (25k articles, teacher labels, all objectives), and IDEA-seg fine-tuning.
Quick start¶
From the repository root, install uv and run the Chonkie baseline on CPU:
cd topic_segmentation
curl -LsSf https://astral.sh/uv/install.sh | sh # install uv (once)
export PATH="$HOME/.local/bin:$PATH" # make uv available in this shell
source scripts/activate.sh # install dependencies and activate
hf download mirth/chonky_modernbert_base_1 --local-dir checkpoints/baselines/chonkie/neural-modernbert-base
python -m topic_segmentation.inference --model-type chonkie_neural --input data/sample_input.jsonl --device cpu
Segment your own articles¶
Replace data/sample_input.jsonl with your input file. Use --model to load a different checkpoint.
ModernBERT (default)¶
Place the trained weights in runs/mind_adaptation/mind_25k_full_news_ts.
python -m topic_segmentation.inference --model-type modernbert_ts --input data/sample_input.jsonl
--window: tokens per window (default: 1024).--stride: tokens between window starts (default: 512).--no-protect-window-end: allow section boundaries at window ends.
Chonkie SemanticChunker¶
Segments by embedding similarity. Weights: checkpoints/baselines/chonkie/semantic-potion-base-32m.
python -m topic_segmentation.inference --model-type chonkie_semantic --input data/sample_input.jsonl
Chonkie NeuralChunker¶
Weights: checkpoints/baselines/chonkie/neural-modernbert-base.
python -m topic_segmentation.inference --model-type chonkie_neural --input data/sample_input.jsonl
Input and output¶
JSONL, one article per line. Example: data/sample_input.jsonl.
{"articleId": "doc_42", "sentences": ["First sentence.", "Second sentence.", "Third sentence."], "labels": [1, 0, 1]}
sentences: list of sentences.articleId: optional; defaults to the line index, starting at 0.labels: optional;1= section ends here,0= section continues.
With labels, evaluation reports precision, recall, F1, Pk, and WindowDiff. The final sentence always gets 1 and is excluded from scoring.
Output fields: articleId, sentences, and predictions (0/1). Use --output to choose the file path. The default is runs/long_document_inference/<model type>_<input>_<time>.jsonl.
Reproduce the thesis experiments¶
Set experiment options and checkpoint paths in configs/. Run from topic_segmentation/:
source scripts/activate.sh # activate the environment
bash scripts/run_ts_pipeline.sh --dry-run # preview all stages
bash scripts/run_ts_pipeline.sh # run all stages
Set GPU IDs and comment out experiments to skip in scripts/run_ts_pipeline.sh. Individual stages are below; --dry-run prints commands and --help lists options.
New runs go to runs/<stage>/reproductions/<time>/, but the configurations keep pointing to the thesis runs: after each stage, set the new checkpoints in the downstream configurations before running the next stage (SETUP.md lists the keys).
Training saves settings and test scores in results.jsonl beside run folders. Inference saves settings and available scores beside the prediction file.
1. Data¶
python scripts/prepare_data.py news # IDEA-seg from saved pages
python scripts/prepare_data.py c4 # paragraph-boundary examples
python scripts/prepare_data.py mimic # clinical evaluation and training splits
MIND data needs the news teacher from stage 3 and is prepared before stage 4.
2. Boundary pre-training¶
python scripts/run_boundary_pretraining.py --gpus 0 1 # learn paragraph boundaries on C4
3. News fine-tuning¶
python scripts/run_news_finetuning.py --experiment backbones --gpu 0 # compare encoders
python scripts/run_news_finetuning.py --experiment boundary_pretraining --gpu 0 # effect of C4 pre-training
python scripts/run_news_finetuning.py --experiment granularity --gpu 0 # paragraphs vs. sentences
python scripts/run_news_finetuning.py --experiment objectives --gpu 0 # training objectives
4. MIND adaptation¶
Set the teacher (teacher_checkpoint in configs/mind_adaptation.json), label MIND, then adapt to the silver labels and fine-tune on IDEA-seg.
python scripts/prepare_data.py mind --gpu 0 # teacher labels and MIND subsets
python scripts/run_mind_adaptation.py --experiment scale --gpu 0 # adaptation set sizes
python scripts/run_mind_adaptation.py --experiment final --gpu 0 # final model
python scripts/run_mind_adaptation.py --experiment objective_matrix --gpu 0 # adaptation and fine-tuning objectives
python scripts/run_mind_adaptation.py --experiment matched_steps --gpu 0 # matched training steps
5. Long-document inference¶
Uses the checkpoints in configs/news_inference.json.
python scripts/run_long_document_inference.py --experiment strides --gpu 0 # stride and window-end protection
python scripts/run_long_document_inference.py --experiment stage_matrix --gpu 0 # models from each training stage
python scripts/run_long_document_inference.py --experiment matched_steps --gpu 0 # matched-step models
python scripts/run_long_document_inference.py --experiment baselines --gpu 0 # Chonkie baselines
6. Clinical transfer¶
Train the news models used for clinical transfer:
python scripts/run_news_finetuning.py --experiment clinical_initialization --gpu 0
python scripts/run_mind_adaptation.py --experiment clinical_initialization --gpu 0
Set their checkpoint paths in configs/mimic_inference.json and configs/mimic_scaling.json, then run:
python scripts/run_clinical_transfer.py zero-shot --gpus 0 # evaluate without clinical fine-tuning
python scripts/run_clinical_transfer.py scaling --gpus 0 1 # vary the clinical training set size
python scripts/run_clinical_transfer.py statistics # results table and paired t-tests
Folder structure¶
topic_segmentation/
├── configs/ # experiment settings and paths
├── scripts/ # commands to run each stage
├── src/topic_segmentation/ # data preparation, models, training, inference, scoring
├── data/ # datasets and sample input
├── checkpoints/ # downloaded model weights
├── cache/ # temporary and cached files
└── runs/ # trained models, predictions, and scores
Data and weights¶
See SETUP.md for downloads and file paths.
Citations¶
@inproceedings{DBLP:conf/emnlp/YuDZLCW23,
author = {Hai Yu and
Chong Deng and
Qinglin Zhang and
Jiaqing Liu and
Qian Chen and
Wen Wang},
editor = {Houda Bouamor and
Juan Pino and
Kalika Bali},
title = {Improving Long Document Topic Segmentation Models With Enhanced Coherence
Modeling},
booktitle = {Proceedings of the 2023 Conference on Empirical Methods in Natural
Language Processing, {EMNLP} 2023, Singapore, December 6-10, 2023},
pages = {5592--5605},
publisher = {Association for Computational Linguistics},
year = {2023},
url = {https://doi.org/10.18653/v1/2023.emnlp-main.341},
doi = {10.18653/V1/2023.EMNLP-MAIN.341},
timestamp = {Wed, 27 Nov 2024 16:25:24 +0100},
biburl = {https://dblp.org/rec/conf/emnlp/YuDZLCW23.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
@inproceedings{DBLP:conf/coling/AbbasiADCF25,
author = {Saeed Abbasi and
Aijun An and
Heidar Davoudi and
Ronald Di Carlantonio and
Gary Farmaner},
editor = {Owen Rambow and
Leo Wanner and
Marianna Apidianaki and
Hend Al{-}Khalifa and
Barbara Di Eugenio and
Steven Schockaert and
Kareem Darwish and
Apoorv Agarwal},
title = {Neural Document Segmentation Using Weighted Sliding Windows with Transformer
Encoders},
booktitle = {Proceedings of the 31st International Conference on Computational
Linguistics, {COLING} 2025 - Industry Track, Abu Dhabi, UAE, January
19-24, 2025},
pages = {807--816},
publisher = {Association for Computational Linguistics},
year = {2025},
url = {https://aclanthology.org/2025.coling-industry.67/},
timestamp = {Tue, 29 Jul 2025 22:41:47 +0200},
biburl = {https://dblp.org/rec/conf/coling/AbbasiADCF25.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
@inproceedings{DBLP:conf/acl/WuQCWQLLXGWZ20,
author = {Fangzhao Wu and
Ying Qiao and
Jiun{-}Hung Chen and
Chuhan Wu and
Tao Qi and
Jianxun Lian and
Danyang Liu and
Xing Xie and
Jianfeng Gao and
Winnie Wu and
Ming Zhou},
editor = {Dan Jurafsky and
Joyce Chai and
Natalie Schluter and
Joel R. Tetreault},
title = {{MIND:} {A} Large-scale Dataset for News Recommendation},
booktitle = {Proceedings of the 58th Annual Meeting of the Association for Computational
Linguistics, {ACL} 2020, Online, July 5-10, 2020},
pages = {3597--3606},
publisher = {Association for Computational Linguistics},
year = {2020},
url = {https://doi.org/10.18653/v1/2020.acl-main.331},
doi = {10.18653/V1/2020.ACL-MAIN.331},
timestamp = {Sat, 15 Aug 2026 09:36:01 +0200},
biburl = {https://dblp.org/rec/conf/acl/WuQCWQLLXGWZ20.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}
@article{DBLP:journals/jmlr/RaffelSRLNMZLL20,
author = {Colin Raffel and
Noam Shazeer and
Adam Roberts and
Katherine Lee and
Sharan Narang and
Michael Matena and
Yanqi Zhou and
Wei Li and
Peter J. Liu},
title = {Exploring the Limits of Transfer Learning with a Unified Text-to-Text
Transformer},
journal = {J. Mach. Learn. Res.},
volume = {21},
pages = {140:1--140:67},
year = {2020},
url = {https://jmlr.org/papers/v21/20-074.html},
timestamp = {Wed, 11 Sep 2024 14:41:27 +0200},
biburl = {https://dblp.org/rec/journals/jmlr/RaffelSRLNMZLL20.bib},
bibsource = {dblp computer science bibliography, https://dblp.org}
}