metrics¶
topic_segmentation.metrics
Score binary topic boundaries: 1 marks the end of a segment.
Pool precision, recall, and F1 over all units. Average Pk and WindowDiff across documents using the convention of Yu et al. (2023). Exclude the document-final boundary.
masses(labels)
¶
Convert binary boundary labels to segment lengths.
Source code in src/topic_segmentation/metrics.py
20 21 22 23 24 25 26 27 28 29 30 | |
window_error(measure, predictions, references)
¶
Return mean window error as 1 - round(mean(1 - error), 4).
Source code in src/topic_segmentation/metrics.py
33 34 35 36 37 38 39 40 41 | |
pk(predictions, references)
¶
Source code in src/topic_segmentation/metrics.py
44 45 | |
windowdiff(predictions, references)
¶
Source code in src/topic_segmentation/metrics.py
48 49 | |
scores(predictions, references)
¶
F1, precision, recall, Pk, and WindowDiff, rounded to 4 decimals.
Source code in src/topic_segmentation/metrics.py
52 53 54 55 56 57 58 | |
trainer_metrics(boundary)
¶
Build a Trainer callback for anchor-view boundary metrics and accuracy.
Source code in src/topic_segmentation/metrics.py
61 62 63 64 65 66 67 68 69 70 71 | |
score_documents(logits, labels, boundary=0)
¶
Score document-level logits against class labels.
Source code in src/topic_segmentation/metrics.py
74 75 76 77 78 | |
score_articles(rows, predictions)
¶
Score labeled articles; return None when no rows have labels.
Source code in src/topic_segmentation/metrics.py
81 82 83 84 85 86 87 88 | |
score_paragraphs(predicted, gold)
¶
Score sentence predictions at paragraph ends using Punkt sentence counts.
Inputs are prediction rows and paragraph-level gold documents. Missing predictions count as continuation (O).
Source code in src/topic_segmentation/metrics.py
91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | |
record(path, run, settings, scores)
¶
Append run settings and scores to a JSONL file.
Source code in src/topic_segmentation/metrics.py
116 117 118 119 120 | |