Environment Setup & Run Commands¶
From the repository root, enter topic_segmentation/ and run the commands below:
cd topic_segmentation
Installation¶
# uv (once)
curl -LsSf https://astral.sh/uv/install.sh | sh
# Create and activate the locked environment, repeat in every new shell
source scripts/activate.sh
# Punkt sentence splitter for IDEA-seg and MIND; activate.sh points NLTK to this folder
python -m nltk.downloader -d data/resources/nltk punkt
Data and weights¶
IDEA-seg, the MIND article bodies, and the trained thesis models are not distributed with the repository. The public models are downloaded below; C4 is downloaded during preparation; MIMIC-IV-Note requires credentialed PhysioNet access.
Models¶
hf download answerdotai/ModernBERT-base --local-dir checkpoints/modernbert_base
hf download google-bert/bert-base-uncased --local-dir checkpoints/bert_base
hf download FacebookAI/roberta-base --local-dir checkpoints/roberta_base
hf download allenai/longformer-base-4096 --local-dir checkpoints/longformer_base
# Chonkie baselines
hf download minishlab/potion-base-32M --local-dir checkpoints/baselines/chonkie/semantic-potion-base-32m
hf download mirth/chonky_modernbert_base_1 --local-dir checkpoints/baselines/chonkie/neural-modernbert-base
Data¶
IDEA-seg¶
The source articles are available in the IDEA dataset. IDEA-seg adds the segmentation labels and splits used in this thesis.
data/news/: train/val/test files, split into paragraphs.data/news_sent/: the same articles, split into sentences.data/news/raw/: saved pages, article list, and source dataset.
Rebuilding from the saved pages writes to runs/news_construction/<time>/; data/news/ and data/news_sent/ keep the thesis copy.
MIND¶
Download the original dataset from the MIND website. The download does not include article bodies; these must be supplied separately (see the official data description).
Put MIND articles with full text in data/mind/full/news.csv. Columns:
news_id,category,subvert,title,abstract,url,entity,ab_entity,body
data/mind/full/: articles and predicted section boundaries.data/mind/prepared/: training splits with 2k, 5k, 10k, and 25k articles.
C4 realnewslike¶
Preparation downloads realnewslike from allenai/c4 on Hugging Face. Tokenized data with paragraph-boundary labels goes in data/c4/tokenized_modernbert/.
MIMIC-IV-Note 2.2¶
Download MIMIC-IV-Note 2.2 from PhysioNet after completing the required training, obtaining credentialed access, and signing the data use agreement.
- Raw notes:
data/mimic/raw/mimic-iv-note-deidentified-free-text-clinical-notes-2.2/. - Prepared evaluation and training data:
data/mimic/.
Check¶
With IDEA-seg and the four backbones in place, check versions, data, and tokenizers offline:
python scripts/checkenv.py
Run Commands¶
Every stage script accepts --dry-run (print the commands only) and --help.
# ── 1. Data ───────────────────────────────────────────────────────────────────
# MIND comes after stage 3: its silver labels need the news teacher (see below)
python scripts/prepare_data.py news c4 mimic
# ── 2. Boundary pre-training on C4 ────────────────────────────────────────────
# Smoke test on 100k examples
python scripts/run_boundary_pretraining.py --smoke-test --gpus 0
# Full run
python scripts/run_boundary_pretraining.py --gpus 0 1
# ── 3. News fine-tuning ───────────────────────────────────────────────────────
# Experiment groups: backbones, boundary_pretraining, granularity, objectives, clinical_initialization
python scripts/run_news_finetuning.py --experiment backbones --gpu 0
python scripts/run_news_finetuning.py --experiment objectives --gpu 0
# One thesis run
python scripts/run_news_finetuning.py --run modernbert_c4_1024_cssl --gpu 0
# Any setting: --backbone bert|roberta|longformer|modernbert|modernbert_c4, --length 512|1024,
# --objective ts|ts_da|cssl|cssl_da|tssp|tssp_da|full_noda|full, --data news|news_sent
python scripts/run_news_finetuning.py --backbone roberta --length 512 --objective ts_da --gpu 0
python scripts/run_news_finetuning.py --backbone modernbert_c4 --length 1024 --objective full --data news_sent --gpu 0
# ── MIND data, once the teacher is set in configs/mind_adaptation.json ─────────
python scripts/prepare_data.py mind --gpu 0
# ── 4. MIND adaptation ────────────────────────────────────────────────────────
# Experiment groups: scale, final, objective_matrix, matched_steps, clinical_initialization
python scripts/run_mind_adaptation.py --experiment final --gpu 0
# One thesis run (reruns the adaptation it starts from)
python scripts/run_mind_adaptation.py --run mind_25k_full_news_ts --gpu 0
# One thesis run, starting from the saved adaptation
python scripts/run_mind_adaptation.py --run mind_25k_full_news_ts --recorded-adaptation --gpu 0
# Any setting: --size 2k|5k|10k|25k, --objective ts|ts_da|full, --matched-steps, --news-objective ts|ts_da|full
python scripts/run_mind_adaptation.py --size 10k --objective full --matched-steps --news-objective ts --gpu 0
# ── 5. Long-document inference ────────────────────────────────────────────────
# Experiment groups: strides, stage_matrix, matched_steps, baselines
python scripts/run_long_document_inference.py --experiment strides --gpu 0
# Any model and window setting
python scripts/run_long_document_inference.py --model runs/mind_adaptation/mind_25k_full_news_ts \
--window 1024 --stride 256 512 768 1024 --no-protect --gpu 0
python scripts/run_long_document_inference.py --model runs/mind_adaptation/mind_25k_full_news_ts \
--window 1024 --stride 512 --protect --gpu 0
# Chonkie baselines
python scripts/run_long_document_inference.py --baseline semantic --gpu 0
python scripts/run_long_document_inference.py --baseline neural --gpu 0
# ── 6. Clinical transfer ──────────────────────────────────────────────────────
# Zero-shot: --models news_pretrain news_ft chonkie_semantic chonkie_neural, --levels block subheading
python scripts/run_clinical_transfer.py zero-shot --gpus 0
python scripts/run_clinical_transfer.py zero-shot --models news_ft --levels block --gpus 0
# Clinical fine-tuning on growing training sets: --initialization base news_pretrain news_ft
python scripts/run_clinical_transfer.py scaling --gpus 0 1
python scripts/run_clinical_transfer.py scaling --initialization base --gpus 0
# Results table and paired t-tests of the runs above -> runs/clinical_transfer/scaling/reproductions/statistics/
python scripts/run_clinical_transfer.py statistics
# ── All stages in thesis order ────────────────────────────────────────────────
# Comment out lines in the script to skip them
bash scripts/run_ts_pipeline.sh --dry-run
bash scripts/run_ts_pipeline.sh
Using reproduced checkpoints¶
Each stage writes new runs to runs/<stage>/reproductions/<time>/, while the configurations point to the thesis runs. The stage scripts and run_ts_pipeline.sh do not pass new weights on: after each stage, update the downstream configuration keys below, then run the next stage. The thesis runs stay in place.
After boundary pre-training¶
Pick a checkpoint of the new run (<run>/epoch-<n> or <run>/checkpoint-<step>) and set it in:
configs/news_finetuning.json:backbones.modernbert_c4configs/mind_adaptation.json:initialization
After news fine-tuning¶
The teacher is modernbert_c4_1024_ts, trained by --experiment boundary_pretraining. Set it, then prepare the MIND data:
configs/mind_adaptation.json:teacher_checkpoint
python scripts/prepare_data.py mind --gpu 0
For long-document inference and clinical transfer, also set:
configs/news_inference.json: themodelentries that nameruns/news_finetuning/modernbert_c4_1024_tsconfigs/mimic_inference.json:models.news_ft.checkpoint(modernbert_c4_1024_full_sentences, trained by--experiment clinical_initialization)configs/mimic_scaling.json:initializations.news_ft
After MIND adaptation¶
Within this stage, --from-adaptation DIR starts the news runs from the adaptations in DIR, as run_ts_pipeline.sh does. Afterwards, set:
configs/news_inference.json: themodelentries that nameruns/mind_adaptation/...configs/mimic_inference.json:models.news_pretrain.checkpoint(mind_25k_cssl_news_ts, trained by--experiment clinical_initialization)configs/mimic_scaling.json:initializations.news_pretrain
Direct inference¶
Pass the new checkpoint instead of editing inference.py:
python -m topic_segmentation.inference --model runs/mind_adaptation/reproductions/<time>/mind_25k_full_news_ts --input data/sample_input.jsonl
Clinical statistics¶
run_clinical_transfer.py statistics reads the reproduced runs of configs/mimic_scaling.json and configs/mimic_inference.json. To rescore the thesis runs:
python -m topic_segmentation.clinical_statistics --scaling runs/clinical_transfer/scaling --zero-shot runs/clinical_transfer/zero_shot