The default bundle and the root lock pin gliner2>=1.3.1,<2 (1.3.2 today). #368 serves the GLiNER2.5-Decide models from the transformers5 bundle, which pins gliner2==2.0.0. It keeps gliner2 1.x everywhere else, because 2.0.0 changes the outputs of fastino/gliner2-base-v1 and fastino/gliner2-large-v1 in two ways. This issue records those changes as input to a later decision on moving those models to 2.x.
How this was measured
- Setup: SIE's
GLiNER2Adapter on fixed inputs, float16 on an L4, transformers 4.57.6, at the pinned revisions. The only difference between runs was gliner2 1.3.2 vs 2.0.0.
- Inputs: six short news and review sentences, plus a ~2,500-word text made of those sentences repeated 40 times.
- Identical under both versions:
- single- and multi-label classification;
- relation extraction with supplied entities;
- NER on one short text;
- an
output_schema with only an enum property;
- GLiGuard (
fastino/gliguard-LLMGuardrails-300M, transformers 5.17).
1. NER returns every mention of a repeated entity
On the long text, entity counts go from 21 to 100 (base-v1) and from 20 to 102 (large-v1). The same 21 and 20 distinct (text, label) pairs come back. Up to 6 mentions of one surface form are now returned where 1.3.2 returned 1. For example, "Airbus" (organization) now appears at offsets 186, 696, 1206, 1716, 2226, …. Scores of the entities present in both runs are identical.
Cause: _format_entity_dict changed its dedup key.
- 1.3.2 deduplicates confidence-and-span results by
text.lower(), keeping only the first mention of each surface form.
- 2.0.0 deduplicates them by
(text.lower(), start, end), keeping every mention.
Arguably 2.0.0 is more correct: each mention comes back with its own offsets. But responses grow, and counts that callers derive from them change.
2. Structured extraction is rescored, and some fields come back null
Raw package outputs for the adapter's output_schema call, with confidence:
| Model |
Text |
Field |
1.3.2 |
2.0.0 |
| base-v1 |
"The iPhone 15 Pro battery drains too fast, but the camera and the titanium frame feel premium." |
person |
"iPhone 15 Pro" (0.72) |
null |
| base-v1 |
"Dr. Priya Raman joined Novartis in Basel as head of oncology research after ten years at Stanford." |
person |
"Dr. Priya Raman" (0.53) |
"Priya Raman" (0.90) |
| large-v1 |
the iPhone text |
sentiment (enum) |
"negative" (0.51) |
null |
| large-v1 |
"I absolutely loved the service at the Hilton in Paris; the staff were friendly and the room was spotless." |
person |
"I" (0.65) |
null |
The adapter rejects a null string property. Under 2.0.0 those requests fail as a whole with GLiNER2 structured extraction property 'person' must be a string, where under 1.3.2 they returned a value. I have not isolated the root cause in the 2.0.0 runtime.
What moving these models to 2.x would take
- Decide whether repeated mentions are the intended output, or whether the adapter should keep 1.x's first-mention behavior.
- Handle
null structured fields: omit the property or return a per-item error, instead of failing the request.
- Re-baseline the
test_all_models entries for both models, and rerun the regression above on the relation and classification paths.
- Then widen the
gliner2 range in packages/sie_server/pyproject.toml and bundles/default.yaml together.
The default bundle and the root lock pin
gliner2>=1.3.1,<2(1.3.2 today). #368 serves the GLiNER2.5-Decide models from the transformers5 bundle, which pinsgliner2==2.0.0. It keeps gliner2 1.x everywhere else, because 2.0.0 changes the outputs offastino/gliner2-base-v1andfastino/gliner2-large-v1in two ways. This issue records those changes as input to a later decision on moving those models to 2.x.How this was measured
GLiNER2Adapteron fixed inputs, float16 on an L4, transformers 4.57.6, at the pinned revisions. The only difference between runs was gliner2 1.3.2 vs 2.0.0.output_schemawith only an enum property;fastino/gliguard-LLMGuardrails-300M, transformers 5.17).1. NER returns every mention of a repeated entity
On the long text, entity counts go from 21 to 100 (base-v1) and from 20 to 102 (large-v1). The same 21 and 20 distinct (text, label) pairs come back. Up to 6 mentions of one surface form are now returned where 1.3.2 returned 1. For example, "Airbus" (organization) now appears at offsets 186, 696, 1206, 1716, 2226, …. Scores of the entities present in both runs are identical.
Cause:
_format_entity_dictchanged its dedup key.text.lower(), keeping only the first mention of each surface form.(text.lower(), start, end), keeping every mention.Arguably 2.0.0 is more correct: each mention comes back with its own offsets. But responses grow, and counts that callers derive from them change.
2. Structured extraction is rescored, and some fields come back null
Raw package outputs for the adapter's
output_schemacall, with confidence:personnullpersonsentiment(enum)nullpersonnullThe adapter rejects a
nullstring property. Under 2.0.0 those requests fail as a whole withGLiNER2 structured extraction property 'person' must be a string, where under 1.3.2 they returned a value. I have not isolated the root cause in the 2.0.0 runtime.What moving these models to 2.x would take
nullstructured fields: omit the property or return a per-item error, instead of failing the request.test_all_modelsentries for both models, and rerun the regression above on the relation and classification paths.gliner2range inpackages/sie_server/pyproject.tomlandbundles/default.yamltogether.