Marine-perimeter classifier (Ifremer positioning study)
Binary classifier: does this scientific publication fall within the marine
perimeter — marine science in the broad sense, as a marine research
institute would understand it? Reads a title and abstract, returns
p_in_perimeter and a binary label at threshold 0.5.
Built by SIRIS Academic for the Ifremer bibliometric positioning study (2026). It is the first level of a two-level pipeline: a recall-first gate in front of four thematic classifiers.
- Private — SIRIS-Lab organisation only. Do not redistribute; whether and how to publish is a SIRIS/Ifremer decision.
- Contact: Pablo Accuosto; Théodore Hervieux / Nicolau Duran Silva (SIRIS).
How to use
Batch scoring of a CSV (recommended — handles Web of Science, Scopus or
OpenAlex exports; auto-detects title/abstract columns including TI/AB):
pip install torch transformers pandas
hf auth login # token with read access to SIRIS-Lab
# get the script (and, to try it out, the bundled 50-row sample):
hf download SIRIS-Lab/ifremer-marine-perimeter mp_classify_csv.py sample_publications.csv --local-dir .
python mp_classify_csv.py sample_publications.csv --out predictions.csv
# then on your own data (WOS tab-separated export: add --sep tab):
python mp_classify_csv.py your_export.csv --out predictions.csv
The output is your CSV plus three columns: p_in_perimeter (probability),
in_perimeter (the label at 0.5), and text_basis — whether each row was
scored on title+abstract or, where the abstract is missing, title_only
(lower reliability; the script also prints how many).
The model itself downloads automatically on first run. Or directly in Python:
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
repo = "SIRIS-Lab/ifremer-marine-perimeter"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForSequenceClassification.from_pretrained(repo).eval()
# one publication = title + blank line + abstract (the training text format)
title = "Larval dispersal and connectivity of flat oyster populations"
abstract = "We combine biophysical modelling and genetics to estimate..."
text = title.strip() + "\n\n" + abstract.strip()
enc = tok(text, truncation=True, max_length=512, return_tensors="pt")
p_in = torch.softmax(model(**enc).logits.float(), -1)[0, 1].item()
in_perimeter = p_in >= 0.5 # 0.5 = the delivered operating point
The input source does not matter — the model reads only title+abstract. WOS, Scopus and OpenAlex records work identically once the two fields are mapped; no adaptation is needed.
Performance (against human-annotated gold)
Evaluated against gold labels annotated by Ifremer's scientific direction (dev n=399; test n=399, spent once, 2026-09-01):
| Split | Accuracy | F1 (IN) | Core stratum acc. |
|---|---|---|---|
| dev | 0.925 | 0.911 | 0.968 (precision 1.0, recall 0.909) |
| test | 0.917 | 0.898 | 0.968 |
The core representative stratum is the closest to a production distribution; the overall figures are on a boundary-enriched draw, so they understate production accuracy. Recall on 178 human-certain positives of a downstream theme: 0.949.
Threshold 0.5 is the delivered operating point: the probability distribution is strongly bimodal (16/399 dev rows in [0.2, 0.8)) and the measured F1 curve is a plateau (≈0.906–0.917) across [0.40, 0.95] — moving the threshold buys nothing. Change it only with a measured reason.
How it was built
BGE-M3 (XLM-RoBERTa-large, 568M, multilingual) fine-tuned with a classification head on 7,316 LLM-labelled rows (silver): two open-weight LLM voters under a validated classification prompt, keeping consensus labels only — including arbiter-resolved disagreement rows was measured and rejected (accuracy 0.925 → 0.890 on gold dev). Human annotation was used only for evaluation, never for training. The perimeter criteria were validated with Ifremer's scientific direction (Jérôme Paillet).
The perimeter is deliberately wide and recall-first: borderline marine papers (coastal engineering, marine socio-economics, land–sea continuum, diadromous species) lean IN; the notable exclusion is strictly freshwater / inland work with no marine link. False negatives are treated as more costly than false positives — downstream steps consume the IN set.
Limitations
- Run it behind a document-type filter. Training and evaluation rows are research publications; editorials, errata, prefaces and other paratext were excluded upstream. On such rows the output is unspecified — filter them out by document type first (as the production pipeline does).
- Title+abstract is the intended input. Title-only rows work but with
lower reliability (
mp_classify_csv.pyflags them). - Indexed scientific literature only; grey literature (expertise reports, technical memos) is out of the training distribution.
- Multilingual via BGE-M3; validated mainly on English and French.
- The IN share on your corpus depends on your corpus: on a marine-enriched candidate pool the model kept ~42% — do not read that as a universal rate.
- Downloads last month
- 77
Model tree for SIRIS-Lab/ifremer-marine-perimeter
Base model
BAAI/bge-m3