Skip to content

Topic Segmentation

Code for news topic segmentation and transfer to clinical notes.

Overview of the topic-segmentation experiments

Read the online documentation for the user guide and API reference.

See SETUP.md for installation, downloads, and parameter examples.

The models mark section boundaries between paragraphs or sentences.

The final model uses ModernBERT-base with C4 boundary pre-training, MIND adaptation (25k articles, teacher labels, all objectives), and IDEA-seg fine-tuning.

Quick start

From the repository root, install uv and run the Chonkie baseline on CPU:

cd topic_segmentation
curl -LsSf https://astral.sh/uv/install.sh | sh # install uv (once)
export PATH="$HOME/.local/bin:$PATH"          # make uv available in this shell
source scripts/activate.sh                   # install dependencies and activate

hf download mirth/chonky_modernbert_base_1 --local-dir checkpoints/baselines/chonkie/neural-modernbert-base
python -m topic_segmentation.inference --model-type chonkie_neural --input data/sample_input.jsonl --device cpu

Segment your own articles

Replace data/sample_input.jsonl with your input file. Use --model to load a different checkpoint.

ModernBERT (default)

Place the trained weights in runs/mind_adaptation/mind_25k_full_news_ts.

python -m topic_segmentation.inference --model-type modernbert_ts --input data/sample_input.jsonl
  • --window: tokens per window (default: 1024).
  • --stride: tokens between window starts (default: 512).
  • --no-protect-window-end: allow section boundaries at window ends.

Chonkie SemanticChunker

Segments by embedding similarity. Weights: checkpoints/baselines/chonkie/semantic-potion-base-32m.

python -m topic_segmentation.inference --model-type chonkie_semantic --input data/sample_input.jsonl

Chonkie NeuralChunker

Weights: checkpoints/baselines/chonkie/neural-modernbert-base.

python -m topic_segmentation.inference --model-type chonkie_neural --input data/sample_input.jsonl

Input and output

JSONL, one article per line. Example: data/sample_input.jsonl.

{"articleId": "doc_42", "sentences": ["First sentence.", "Second sentence.", "Third sentence."], "labels": [1, 0, 1]}
  • sentences: list of sentences.
  • articleId: optional; defaults to the line index, starting at 0.
  • labels: optional; 1 = section ends here, 0 = section continues.

With labels, evaluation reports precision, recall, F1, Pk, and WindowDiff. The final sentence always gets 1 and is excluded from scoring.

Output fields: articleId, sentences, and predictions (0/1). Use --output to choose the file path. The default is runs/long_document_inference/<model type>_<input>_<time>.jsonl.

Reproduce the thesis experiments

Set experiment options and checkpoint paths in configs/. Run from topic_segmentation/:

source scripts/activate.sh                # activate the environment
bash scripts/run_ts_pipeline.sh --dry-run # preview all stages
bash scripts/run_ts_pipeline.sh           # run all stages

Set GPU IDs and comment out experiments to skip in scripts/run_ts_pipeline.sh. Individual stages are below; --dry-run prints commands and --help lists options.

New runs go to runs/<stage>/reproductions/<time>/, but the configurations keep pointing to the thesis runs: after each stage, set the new checkpoints in the downstream configurations before running the next stage (SETUP.md lists the keys).

Training saves settings and test scores in results.jsonl beside run folders. Inference saves settings and available scores beside the prediction file.

1. Data

python scripts/prepare_data.py news  # IDEA-seg from saved pages
python scripts/prepare_data.py c4    # paragraph-boundary examples
python scripts/prepare_data.py mimic # clinical evaluation and training splits

MIND data needs the news teacher from stage 3 and is prepared before stage 4.

2. Boundary pre-training

python scripts/run_boundary_pretraining.py --gpus 0 1 # learn paragraph boundaries on C4

3. News fine-tuning

python scripts/run_news_finetuning.py --experiment backbones --gpu 0            # compare encoders
python scripts/run_news_finetuning.py --experiment boundary_pretraining --gpu 0 # effect of C4 pre-training
python scripts/run_news_finetuning.py --experiment granularity --gpu 0          # paragraphs vs. sentences
python scripts/run_news_finetuning.py --experiment objectives --gpu 0           # training objectives

4. MIND adaptation

Set the teacher (teacher_checkpoint in configs/mind_adaptation.json), label MIND, then adapt to the silver labels and fine-tune on IDEA-seg.

python scripts/prepare_data.py mind --gpu 0                                 # teacher labels and MIND subsets
python scripts/run_mind_adaptation.py --experiment scale --gpu 0            # adaptation set sizes
python scripts/run_mind_adaptation.py --experiment final --gpu 0            # final model
python scripts/run_mind_adaptation.py --experiment objective_matrix --gpu 0 # adaptation and fine-tuning objectives
python scripts/run_mind_adaptation.py --experiment matched_steps --gpu 0    # matched training steps

5. Long-document inference

Uses the checkpoints in configs/news_inference.json.

python scripts/run_long_document_inference.py --experiment strides --gpu 0       # stride and window-end protection
python scripts/run_long_document_inference.py --experiment stage_matrix --gpu 0  # models from each training stage
python scripts/run_long_document_inference.py --experiment matched_steps --gpu 0 # matched-step models
python scripts/run_long_document_inference.py --experiment baselines --gpu 0     # Chonkie baselines

6. Clinical transfer

Train the news models used for clinical transfer:

python scripts/run_news_finetuning.py --experiment clinical_initialization --gpu 0
python scripts/run_mind_adaptation.py --experiment clinical_initialization --gpu 0

Set their checkpoint paths in configs/mimic_inference.json and configs/mimic_scaling.json, then run:

python scripts/run_clinical_transfer.py zero-shot --gpus 0 # evaluate without clinical fine-tuning
python scripts/run_clinical_transfer.py scaling --gpus 0 1 # vary the clinical training set size
python scripts/run_clinical_transfer.py statistics        # results table and paired t-tests

Folder structure

topic_segmentation/
├── configs/                 # experiment settings and paths
├── scripts/                 # commands to run each stage
├── src/topic_segmentation/  # data preparation, models, training, inference, scoring
├── data/                    # datasets and sample input
├── checkpoints/             # downloaded model weights
├── cache/                   # temporary and cached files
└── runs/                    # trained models, predictions, and scores

Data and weights

See SETUP.md for downloads and file paths.

Citations

@inproceedings{DBLP:conf/emnlp/YuDZLCW23,
  author       = {Hai Yu and
                  Chong Deng and
                  Qinglin Zhang and
                  Jiaqing Liu and
                  Qian Chen and
                  Wen Wang},
  editor       = {Houda Bouamor and
                  Juan Pino and
                  Kalika Bali},
  title        = {Improving Long Document Topic Segmentation Models With Enhanced Coherence
                  Modeling},
  booktitle    = {Proceedings of the 2023 Conference on Empirical Methods in Natural
                  Language Processing, {EMNLP} 2023, Singapore, December 6-10, 2023},
  pages        = {5592--5605},
  publisher    = {Association for Computational Linguistics},
  year         = {2023},
  url          = {https://doi.org/10.18653/v1/2023.emnlp-main.341},
  doi          = {10.18653/V1/2023.EMNLP-MAIN.341},
  timestamp    = {Wed, 27 Nov 2024 16:25:24 +0100},
  biburl       = {https://dblp.org/rec/conf/emnlp/YuDZLCW23.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org}
}

@inproceedings{DBLP:conf/coling/AbbasiADCF25,
  author       = {Saeed Abbasi and
                  Aijun An and
                  Heidar Davoudi and
                  Ronald Di Carlantonio and
                  Gary Farmaner},
  editor       = {Owen Rambow and
                  Leo Wanner and
                  Marianna Apidianaki and
                  Hend Al{-}Khalifa and
                  Barbara Di Eugenio and
                  Steven Schockaert and
                  Kareem Darwish and
                  Apoorv Agarwal},
  title        = {Neural Document Segmentation Using Weighted Sliding Windows with Transformer
                  Encoders},
  booktitle    = {Proceedings of the 31st International Conference on Computational
                  Linguistics, {COLING} 2025 - Industry Track, Abu Dhabi, UAE, January
                  19-24, 2025},
  pages        = {807--816},
  publisher    = {Association for Computational Linguistics},
  year         = {2025},
  url          = {https://aclanthology.org/2025.coling-industry.67/},
  timestamp    = {Tue, 29 Jul 2025 22:41:47 +0200},
  biburl       = {https://dblp.org/rec/conf/coling/AbbasiADCF25.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org}
}

@inproceedings{DBLP:conf/acl/WuQCWQLLXGWZ20,
  author       = {Fangzhao Wu and
                  Ying Qiao and
                  Jiun{-}Hung Chen and
                  Chuhan Wu and
                  Tao Qi and
                  Jianxun Lian and
                  Danyang Liu and
                  Xing Xie and
                  Jianfeng Gao and
                  Winnie Wu and
                  Ming Zhou},
  editor       = {Dan Jurafsky and
                  Joyce Chai and
                  Natalie Schluter and
                  Joel R. Tetreault},
  title        = {{MIND:} {A} Large-scale Dataset for News Recommendation},
  booktitle    = {Proceedings of the 58th Annual Meeting of the Association for Computational
                  Linguistics, {ACL} 2020, Online, July 5-10, 2020},
  pages        = {3597--3606},
  publisher    = {Association for Computational Linguistics},
  year         = {2020},
  url          = {https://doi.org/10.18653/v1/2020.acl-main.331},
  doi          = {10.18653/V1/2020.ACL-MAIN.331},
  timestamp    = {Sat, 15 Aug 2026 09:36:01 +0200},
  biburl       = {https://dblp.org/rec/conf/acl/WuQCWQLLXGWZ20.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org}
}

@article{DBLP:journals/jmlr/RaffelSRLNMZLL20,
  author       = {Colin Raffel and
                  Noam Shazeer and
                  Adam Roberts and
                  Katherine Lee and
                  Sharan Narang and
                  Michael Matena and
                  Yanqi Zhou and
                  Wei Li and
                  Peter J. Liu},
  title        = {Exploring the Limits of Transfer Learning with a Unified Text-to-Text
                  Transformer},
  journal      = {J. Mach. Learn. Res.},
  volume       = {21},
  pages        = {140:1--140:67},
  year         = {2020},
  url          = {https://jmlr.org/papers/v21/20-074.html},
  timestamp    = {Wed, 11 Sep 2024 14:41:27 +0200},
  biburl       = {https://dblp.org/rec/journals/jmlr/RaffelSRLNMZLL20.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org}
}