Skip to content

gliner2 2.x changes gliner2-base-v1 and gliner2-large-v1 outputs (repeated NER mentions, structured extraction) #369

Description

@svonava

The default bundle and the root lock pin gliner2>=1.3.1,<2 (1.3.2 today). #368 serves the GLiNER2.5-Decide models from the transformers5 bundle, which pins gliner2==2.0.0. It keeps gliner2 1.x everywhere else, because 2.0.0 changes the outputs of fastino/gliner2-base-v1 and fastino/gliner2-large-v1 in two ways. This issue records those changes as input to a later decision on moving those models to 2.x.

How this was measured

  • Setup: SIE's GLiNER2Adapter on fixed inputs, float16 on an L4, transformers 4.57.6, at the pinned revisions. The only difference between runs was gliner2 1.3.2 vs 2.0.0.
  • Inputs: six short news and review sentences, plus a ~2,500-word text made of those sentences repeated 40 times.
  • Identical under both versions:
    • single- and multi-label classification;
    • relation extraction with supplied entities;
    • NER on one short text;
    • an output_schema with only an enum property;
    • GLiGuard (fastino/gliguard-LLMGuardrails-300M, transformers 5.17).

1. NER returns every mention of a repeated entity

On the long text, entity counts go from 21 to 100 (base-v1) and from 20 to 102 (large-v1). The same 21 and 20 distinct (text, label) pairs come back. Up to 6 mentions of one surface form are now returned where 1.3.2 returned 1. For example, "Airbus" (organization) now appears at offsets 186, 696, 1206, 1716, 2226, …. Scores of the entities present in both runs are identical.

Cause: _format_entity_dict changed its dedup key.

  • 1.3.2 deduplicates confidence-and-span results by text.lower(), keeping only the first mention of each surface form.
  • 2.0.0 deduplicates them by (text.lower(), start, end), keeping every mention.

Arguably 2.0.0 is more correct: each mention comes back with its own offsets. But responses grow, and counts that callers derive from them change.

2. Structured extraction is rescored, and some fields come back null

Raw package outputs for the adapter's output_schema call, with confidence:

Model Text Field 1.3.2 2.0.0
base-v1 "The iPhone 15 Pro battery drains too fast, but the camera and the titanium frame feel premium." person "iPhone 15 Pro" (0.72) null
base-v1 "Dr. Priya Raman joined Novartis in Basel as head of oncology research after ten years at Stanford." person "Dr. Priya Raman" (0.53) "Priya Raman" (0.90)
large-v1 the iPhone text sentiment (enum) "negative" (0.51) null
large-v1 "I absolutely loved the service at the Hilton in Paris; the staff were friendly and the room was spotless." person "I" (0.65) null

The adapter rejects a null string property. Under 2.0.0 those requests fail as a whole with GLiNER2 structured extraction property 'person' must be a string, where under 1.3.2 they returned a value. I have not isolated the root cause in the 2.0.0 runtime.

What moving these models to 2.x would take

  • Decide whether repeated mentions are the intended output, or whether the adapter should keep 1.x's first-mention behavior.
  • Handle null structured fields: omit the property or return a per-item error, instead of failing the request.
  • Re-baseline the test_all_models entries for both models, and rerun the regression above on the relation and classification paths.
  • Then widen the gliner2 range in packages/sie_server/pyproject.toml and bundles/default.yaml together.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions