Skip to content

Environment Setup & Run Commands

From the repository root, enter topic_segmentation/ and run the commands below:

cd topic_segmentation

Installation

# uv (once)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Create and activate the locked environment, repeat in every new shell
source scripts/activate.sh

# Punkt sentence splitter for IDEA-seg and MIND; activate.sh points NLTK to this folder
python -m nltk.downloader -d data/resources/nltk punkt

Data and weights

IDEA-seg, the MIND article bodies, and the trained thesis models are not distributed with the repository. The public models are downloaded below; C4 is downloaded during preparation; MIMIC-IV-Note requires credentialed PhysioNet access.

Models

hf download answerdotai/ModernBERT-base --local-dir checkpoints/modernbert_base
hf download google-bert/bert-base-uncased --local-dir checkpoints/bert_base
hf download FacebookAI/roberta-base --local-dir checkpoints/roberta_base
hf download allenai/longformer-base-4096 --local-dir checkpoints/longformer_base

# Chonkie baselines
hf download minishlab/potion-base-32M --local-dir checkpoints/baselines/chonkie/semantic-potion-base-32m
hf download mirth/chonky_modernbert_base_1 --local-dir checkpoints/baselines/chonkie/neural-modernbert-base

Data

IDEA-seg

The source articles are available in the IDEA dataset. IDEA-seg adds the segmentation labels and splits used in this thesis.

  • data/news/: train/val/test files, split into paragraphs.
  • data/news_sent/: the same articles, split into sentences.
  • data/news/raw/: saved pages, article list, and source dataset.

Rebuilding from the saved pages writes to runs/news_construction/<time>/; data/news/ and data/news_sent/ keep the thesis copy.

MIND

Download the original dataset from the MIND website. The download does not include article bodies; these must be supplied separately (see the official data description).

Put MIND articles with full text in data/mind/full/news.csv. Columns:

news_id,category,subvert,title,abstract,url,entity,ab_entity,body
  • data/mind/full/: articles and predicted section boundaries.
  • data/mind/prepared/: training splits with 2k, 5k, 10k, and 25k articles.

C4 realnewslike

Preparation downloads realnewslike from allenai/c4 on Hugging Face. Tokenized data with paragraph-boundary labels goes in data/c4/tokenized_modernbert/.

MIMIC-IV-Note 2.2

Download MIMIC-IV-Note 2.2 from PhysioNet after completing the required training, obtaining credentialed access, and signing the data use agreement.

  • Raw notes: data/mimic/raw/mimic-iv-note-deidentified-free-text-clinical-notes-2.2/.
  • Prepared evaluation and training data: data/mimic/.

Check

With IDEA-seg and the four backbones in place, check versions, data, and tokenizers offline:

python scripts/checkenv.py

Run Commands

Every stage script accepts --dry-run (print the commands only) and --help.

# ── 1. Data ───────────────────────────────────────────────────────────────────
# MIND comes after stage 3: its silver labels need the news teacher (see below)
python scripts/prepare_data.py news c4 mimic

# ── 2. Boundary pre-training on C4 ────────────────────────────────────────────
# Smoke test on 100k examples
python scripts/run_boundary_pretraining.py --smoke-test --gpus 0

# Full run
python scripts/run_boundary_pretraining.py --gpus 0 1

# ── 3. News fine-tuning ───────────────────────────────────────────────────────
# Experiment groups: backbones, boundary_pretraining, granularity, objectives, clinical_initialization
python scripts/run_news_finetuning.py --experiment backbones --gpu 0
python scripts/run_news_finetuning.py --experiment objectives --gpu 0

# One thesis run
python scripts/run_news_finetuning.py --run modernbert_c4_1024_cssl --gpu 0

# Any setting: --backbone bert|roberta|longformer|modernbert|modernbert_c4, --length 512|1024,
# --objective ts|ts_da|cssl|cssl_da|tssp|tssp_da|full_noda|full, --data news|news_sent
python scripts/run_news_finetuning.py --backbone roberta --length 512 --objective ts_da --gpu 0
python scripts/run_news_finetuning.py --backbone modernbert_c4 --length 1024 --objective full --data news_sent --gpu 0

# ── MIND data, once the teacher is set in configs/mind_adaptation.json ─────────
python scripts/prepare_data.py mind --gpu 0

# ── 4. MIND adaptation ────────────────────────────────────────────────────────
# Experiment groups: scale, final, objective_matrix, matched_steps, clinical_initialization
python scripts/run_mind_adaptation.py --experiment final --gpu 0

# One thesis run (reruns the adaptation it starts from)
python scripts/run_mind_adaptation.py --run mind_25k_full_news_ts --gpu 0

# One thesis run, starting from the saved adaptation
python scripts/run_mind_adaptation.py --run mind_25k_full_news_ts --recorded-adaptation --gpu 0

# Any setting: --size 2k|5k|10k|25k, --objective ts|ts_da|full, --matched-steps, --news-objective ts|ts_da|full
python scripts/run_mind_adaptation.py --size 10k --objective full --matched-steps --news-objective ts --gpu 0

# ── 5. Long-document inference ────────────────────────────────────────────────
# Experiment groups: strides, stage_matrix, matched_steps, baselines
python scripts/run_long_document_inference.py --experiment strides --gpu 0

# Any model and window setting
python scripts/run_long_document_inference.py --model runs/mind_adaptation/mind_25k_full_news_ts \
    --window 1024 --stride 256 512 768 1024 --no-protect --gpu 0
python scripts/run_long_document_inference.py --model runs/mind_adaptation/mind_25k_full_news_ts \
    --window 1024 --stride 512 --protect --gpu 0

# Chonkie baselines
python scripts/run_long_document_inference.py --baseline semantic --gpu 0
python scripts/run_long_document_inference.py --baseline neural --gpu 0

# ── 6. Clinical transfer ──────────────────────────────────────────────────────
# Zero-shot: --models news_pretrain news_ft chonkie_semantic chonkie_neural, --levels block subheading
python scripts/run_clinical_transfer.py zero-shot --gpus 0
python scripts/run_clinical_transfer.py zero-shot --models news_ft --levels block --gpus 0

# Clinical fine-tuning on growing training sets: --initialization base news_pretrain news_ft
python scripts/run_clinical_transfer.py scaling --gpus 0 1
python scripts/run_clinical_transfer.py scaling --initialization base --gpus 0

# Results table and paired t-tests of the runs above -> runs/clinical_transfer/scaling/reproductions/statistics/
python scripts/run_clinical_transfer.py statistics

# ── All stages in thesis order ────────────────────────────────────────────────
# Comment out lines in the script to skip them
bash scripts/run_ts_pipeline.sh --dry-run
bash scripts/run_ts_pipeline.sh

Using reproduced checkpoints

Each stage writes new runs to runs/<stage>/reproductions/<time>/, while the configurations point to the thesis runs. The stage scripts and run_ts_pipeline.sh do not pass new weights on: after each stage, update the downstream configuration keys below, then run the next stage. The thesis runs stay in place.

After boundary pre-training

Pick a checkpoint of the new run (<run>/epoch-<n> or <run>/checkpoint-<step>) and set it in:

  • configs/news_finetuning.json: backbones.modernbert_c4
  • configs/mind_adaptation.json: initialization

After news fine-tuning

The teacher is modernbert_c4_1024_ts, trained by --experiment boundary_pretraining. Set it, then prepare the MIND data:

  • configs/mind_adaptation.json: teacher_checkpoint
python scripts/prepare_data.py mind --gpu 0

For long-document inference and clinical transfer, also set:

  • configs/news_inference.json: the model entries that name runs/news_finetuning/modernbert_c4_1024_ts
  • configs/mimic_inference.json: models.news_ft.checkpoint (modernbert_c4_1024_full_sentences, trained by --experiment clinical_initialization)
  • configs/mimic_scaling.json: initializations.news_ft

After MIND adaptation

Within this stage, --from-adaptation DIR starts the news runs from the adaptations in DIR, as run_ts_pipeline.sh does. Afterwards, set:

  • configs/news_inference.json: the model entries that name runs/mind_adaptation/...
  • configs/mimic_inference.json: models.news_pretrain.checkpoint (mind_25k_cssl_news_ts, trained by --experiment clinical_initialization)
  • configs/mimic_scaling.json: initializations.news_pretrain

Direct inference

Pass the new checkpoint instead of editing inference.py:

python -m topic_segmentation.inference --model runs/mind_adaptation/reproductions/<time>/mind_25k_full_news_ts --input data/sample_input.jsonl

Clinical statistics

run_clinical_transfer.py statistics reads the reproduced runs of configs/mimic_scaling.json and configs/mimic_inference.json. To rescore the thesis runs:

python -m topic_segmentation.clinical_statistics --scaling runs/clinical_transfer/scaling --zero-shot runs/clinical_transfer/zero_shot