Skip to content

Panoptic segmentation

Panoptic segmentation labels every pixel with a category and, for countable "things", an instance — one map per image that covers people and cars as instances and sky and road as regions. Its metric is panoptic quality (PQ) (Kirillov et al., CVPR 2019), and the reference implementation is panopticapi. hotcoco computes PQ with the same rules and the same numbers, from the same files or from annotations that carry masks, and reports it in the same EvalReport shape as detection.

What PQ measures

Ground-truth and predicted segments of the same category are matched when their IoU is above 0.5. The threshold makes the match unique in both directions, so no assignment solver is involved. From the matched pairs (TP), the unmatched ground truth (FN), and the unmatched predictions (FP):

PQ = Σ IoU / (TP + ½ FP + ½ FN)      segmentation and recognition together
SQ = Σ IoU / TP                      mean IoU of the matched pairs
RQ = TP / (TP + ½ FP + ½ FN)         an F1 over segments

so PQ = SQ × RQ. Each is computed per category, then averaged over the categories that have at least one segment on either side — over all of them, over things, and over stuff. Two rules keep the protocol fair to the model:

  • a prediction more than half covered by void (unlabeled pixels) or by a crowd region of its own category is ignored rather than counted as a false positive, and void pixels are left out of the IoU denominator;
  • a crowd segment is never matched and never a miss.

Evaluate the COCO panoptic format

The COCO panoptic release is a JSON file plus a folder of PNG files, one per image, where each pixel's color encodes its segment id. Predictions take the same shape. Point PanopticEval at the two JSON files; the PNG folders default to each path without its .json, which is where the COCO release and panopticapi put them:

from hotcoco import panoptic

ev = panoptic.PanopticEval("panoptic_val2017.json", "predictions.json")
ev.run()
          |    PQ     SQ     RQ     N
--------------------------------------
All       |  62.5   93.5   66.8   133
Things    |  60.4   93.2   64.8    80
Stuff     |  65.6   93.9   69.9    53

Pass gt_folder= and pred_folder= when the PNG files live elsewhere. The predictions' categories are ignored; the ground truth's are scored, as in panopticapi. evaluate() and summarize() are the two halves of run() when you want the numbers without the table.

Drop-in for pq_compute

Code that calls panopticapi's one function changes one import:

from hotcoco.panoptic import pq_compute   # was: from panopticapi.evaluation import pq_compute

res = pq_compute("panoptic_val2017.json", "predictions.json")
res["All"]["pq"], res["Things"]["pq"], res["Stuff"]["pq"]

Same arguments, same return shape, scores as fractions. The one difference is in per_class: a category with no segment on either side reports -1.0, hotcoco's "not computed" marker, where panopticapi prints 0.0 and then leaves it out of the average. The tp, fp, fn, and iou counts beside each class are additions.

Evaluate without PNG files

The PNG round trip is a packaging choice, not part of the metric. A detection-style dataset whose annotations carry masks — RLE or polygon, one annotation per segment, iscrowd where it applies — evaluates directly:

from hotcoco import COCO, panoptic

gt = COCO("instances_panoptic.json")          # one annotation per segment, with masks
pred = gt.load_res("segments.json")           # the same shape; scores are ignored

ev = panoptic.PanopticEval(gt, pred)
ev.run()

Masks are rasterized onto the image's height × width, later annotations over earlier ones where they overlap, and the pixels found are the segment — a stale area field cannot move an IoU on this path. The two inputs can be mixed: ground truth from PNG files, predictions as masks, or the reverse. On the same pixels, both paths give identical numbers.

Categories need isthing (1 or true for things, 0 or false for stuff) for the things and stuff splits. Without it a category is scored in All and in neither split, and provenance reports "extension" because panopticapi would not have evaluated the file at all — see provenance.

Results and the report

results() is panopticapi's dict with the counts added; report() is the family-neutral form:

r = ev.report()
r["metrics"]["PQ"], r["metrics"]["PQ_th"], r["metrics"]["PQ_st"]
r["per_class"]["person"]           # {"PQ": ..., "SQ": ..., "RQ": ...}
r["per_group"]["stuff"]["n"]       # categories averaged into the stuff split
r["provenance"]                    # "parity_verified"

ev.stats holds the nine headline values in panoptic.METRIC_NAMES order — PQ, SQ, RQ, PQ_th, SQ_th, RQ_th, PQ_st, SQ_st, RQ_st, fractions in [0, 1] — which is what detection frameworks log, usually multiplied by 100.

From the CLI

coco panoptic eval --gt panoptic_val2017.json --pred predictions.json
coco panoptic eval --gt gt.json --pred pred.json --gt-folder gt_png/ --pred-folder pred_png/ --json

The Rust binary has the same subcommand: coco-eval panoptic --gt ... --pred .... Both are in the CLI reference.

What panopticapi rejects, hotcoco rejects

A predicted segment listed in the JSON but absent from its PNG, a PNG id missing from the JSON, a prediction with a category the ground truth does not define, and a ground-truth image with no prediction are all errors, raised from evaluate() with the image id. A ground-truth PNG id with no JSON entry is accepted and treated as neither a segment nor void, as the reference does.

Where panopticapi would divide by zero — a things or stuff split with no evaluable category — hotcoco reports -1.0 with n = 0 instead of failing.

Verification

PQ, SQ, and RQ match panopticapi on COCO panoptic val2017 to the last printed digit, with every per-category count identical — the figures are in Panoptic parity. The same comparison runs in CI on synthetic label maps that exercise every rule above.