Open-world counters count whatever you name, but only what they can see: put a box over part of a tray and they report the visible objects and stop. The reason is architectural. Under an occluder, the backbone encodes the occluder. CountOCC reconstructs the missing object features at every level of the feature pyramid, from the visible context and the text and exemplar prompt, and trains the occluded view to attend where an unoccluded view would.
Figure 1. The occlusion challenge. (a) Twelve donuts, all visible. (b) An occluder hides two. (c) A state-of-the-art counter finds the ten it can see and misses the rest. (d) CountOCC places density inside the occluder and recovers both hidden donuts, for a total of 12.
Counters count what they see
Open-world counting went from one model per category to models told what to count at test time, by visual exemplars (CounTR, LOCA), by text (CLIP-Count, CounTX), or by both (CountGD, the current state of the art). All of them assume the targets are mostly visible. Humans do not: we infer the donuts under the napkin from the rows around it.
Amodal counting asks for the total, visible or not. Given an image \(X_I\), exemplar boxes \(\mathcal{B}\), a text prompt \(t\) and an occlusion mask \(M_o\), the model must return
The benchmark CAPTURe tests this on regular patterns, where the hidden objects can be read off the visible ones. Real scenes are often unstructured, and nothing outside the occluder says what is under it. To cover both, the paper builds occlusion-augmented versions of two standard benchmarks, keeping their splits and annotations: FSC-147-OCC (147 categories, with disjoint training, validation and test classes) and CARPK-OCC (parking lots seen from a drone). Black rectangles, at most 256 px on a side, are centred on annotated objects until 25–35% of the instances in each image are covered.
Examples from the two benchmarks FSC-147-OCC and CARPK-OCC
Reconstruct the features, then check the attention
CountOCC builds on CountGD and adds two forms of supervision. One works in feature space: rebuild what the backbone would have seen without the occluder. The other works in attention space: make the occluded view look where the clean view looks.
Figure 2. The CountOCC architecture. At each pyramid level, the Feature Reconstruction Module (FRM) replaces corrupted occluded tokens with reconstructed ones. Visual Equivalence (VisEQ) aligns gradient-based attention maps of a teacher that sees the clean image and a student that sees the occluded one. The reconstructed features flow through CountGD's feature enhancer and cross-modality decoder, which count the visible and the occluded instances.
Feature Reconstruction Module
The Swin backbone produces features at three pyramid levels (256, 512 and 1024 channels). At each level \(\ell\), tokens outside the mask are kept as visible context \(Z_{\text{vis}}\), and every occluded position starts from a learned mask embedding. The queries first attend to each other, then to the visible context, then to the fused text and exemplar tokens \(Z_{v,t}\), which keeps the reconstruction on the right category:
The reconstructed tokens replace the occluded positions, so the decoder sees a complete feature pyramid. They are supervised by a frozen teacher that sees the clean image: on occluded positions only, a Charbonnier term, a cosine term and an \(\ell_2\) term pull the student's reconstruction \(\hat Z_S\) toward the teacher's features \(\hat Z_T\), with \(\Delta = \hat Z_S - \hat Z_T\) and weights \((\lambda_{\text{charb}}, \lambda_{\cos}, \lambda_{\ell_2}) = (0.3, 0.5, 1.0)\):
Visual Equivalence
Good features are not enough if the model looks in the wrong place. Inspired by SelfEQ, CountOCC computes a language-conditioned Grad-CAM map \(G\) for each view: the gradient of the matching score (the mean of the top 900 query confidences) weights the channels of each pyramid level, and the levels are blended by their gradient energy. The student's map \(G_S\) on the occluded image is pulled toward the teacher's \(G_T\) on the clean one:
The second term looks only at the region of interest where either map is confident (\(G_T + G_S \ge \tau\)): it asks for high and uniform activation there, which rules out the trivial solution of two maps that agree by both being flat.
Figure 3. The two modules. (a) FRM: learnable queries at occluded positions attend to each other, to the visible tokens and to the fused prompt, and an MLP turns them into reconstructed features. (b) VisEQ: the teacher's attention on the clean image and the student's on the occluded image are aligned by \(\mathcal{L}_{\text{sim}}\) and made confident by \(\mathcal{L}_{\text{cst}}\).
- Stage 1 · reconstructionTrain FRM alone on FSC-147 training images with occluders drawn on the fly: 128–256 px black rectangles anchored on objects until 15–50% of them are covered. The Swin-B image and BERT text encoders stay frozen.
- Stage 2 · attentionStarting from the best FRM checkpoint, train FRM and VisEQ together with the counting loss.
- InferenceOnly the student runs. The teacher and VisEQ exist only in training, so CountOCC adds no cost beyond the reconstruction module.
Lower error on four occluded benchmarks
Baselines run from their official checkpoints under their own protocols, with the same occluded inputs as CountOCC. One model, trained only on FSC-147, is evaluated everywhere.
The FSC-147 validation and test splits with black rectangles laid over 25–35% of the annotated objects in every image. The test split's 29 categories never appear in training.
11.42 test MAE against 14.42 for CountGD, the strongest open-world counter: −20.8% on test and −26.7% on validation. RMSE falls further, by 54.7% on test, so the worst failures shrink the most.
Text-only counters suffer most under occlusion; CountOCC roughly halves the error of CounTX and CLIP-Count.
MAE and RMSE · FSC-147-OCC 6 methods
| Metric | CLIP-Count | CounTX | CounTR | LOCA | CountGD | CountOCC |
|---|---|---|---|---|---|---|
| Validation MAE | 26.31 | 24.81 | 23.14 | 17.13 | 15.83 | 11.60 |
| Validation RMSE | 80.45 | 75.58 | 66.78 | 44.25 | 54.38 | 35.40 |
| Test MAE | 23.90 | 23.04 | 22.25 | 16.77 | 14.42 | 11.42 |
| Test RMSE | 108.57 | 113.83 | 104.75 | 78.41 | 85.40 | 38.68 |
Lower is better. Bold marks the best in each row and underline the runner-up. Targets are specified by text (CLIP-Count, CounTX), visual exemplars (CounTR, LOCA) or both (CountGD, CountOCC).
Drone images of parking lots with the same occlusion scheme, evaluated zero-shot: CountOCC never sees a CARPK image in training.
4.65 MAE against 9.28 for CountGD: the error halves (−49.9%), and RMSE falls by 47.6%.
CounTR was fine-tuned on the original CARPK and still ends at 14.99.
MAE and RMSE · CARPK-OCC 6 methods
| Metric | CLIP-Count | CounTX | CounTR | LOCA | CountGD | CountOCC |
|---|---|---|---|---|---|---|
| Test MAE | 17.43 | 12.58 | 14.99 | 22.02 | 9.28 | 4.65 |
| Test RMSE | 20.74 | 15.40 | 16.84 | 24.55 | 11.27 | 5.91 |
Lower is better. Bold marks the best in each row and underline the runner-up. Two annotated boxes per image serve as visual exemplars, and the text prompt is “car”.
924 FSC-147 images in which people placed occluders over regular, repeated patterns, so hidden objects can be inferred from the pattern.
10.66 MAE against 14.97 for CountGD, a 28.8% reduction, with the same model and no further training.
RMSE barely moves (41.62 to 41.31), so a few large errors remain on this benchmark.
MAE and RMSE · CAPTURe-Real 2 methods
| Metric | CountGD | CountOCC |
|---|---|---|
| MAE | 14.97 | 10.66 |
| RMSE | 41.62 | 41.31 |
Lower is better. Bold marks the best in each row and underline the runner-up. Only CountGD is reported as a baseline on CAPTURe-Real.
Real crowds, where people hide each other. CrowdHuman's annotated overlap regions serve as the occlusion masks.
8.24 MAE against 9.97 for CountGD (−17.4%), and RMSE drops from 24.46 to 17.87 (−26.9%).
Here occlusion is natural rather than a synthetic black mask.
MAE and RMSE · CrowdHuman 2 methods
| Metric | CountGD | CountOCC |
|---|---|---|
| MAE | 9.97 | 8.24 |
| RMSE | 24.46 | 17.87 |
Lower is better. Bold marks the best in each row and underline the runner-up. The ground-truth total is the number of annotated people in each image.
The gain is on the hidden objects
Splitting the error by instance type shows where the improvement comes from. Counts inside the mask and outside it are scored separately.
On occluded instances the error falls from 18.16 to 5.80, to about a third; on validation, from 17.46 to 5.61. On visible instances it barely moves: 9.74 to 8.52 on test and 8.05 to 8.04 on validation. CountOCC has learned to reason about what is hidden, not just to count the visible objects better.
Visible-instance RMSE is slightly higher than CountGD's on both splits (30.03 vs 27.04, 55.28 vs 53.94).
Figure 4. Density maps on occluded FSC-147 scenes. Columns: CLIP-Count, CounTX, CounTR, LOCA, CountGD and CountOCC. The baselines leave the occluded regions empty and undercount; CountOCC places density inside the masks.
What each piece does
Reconstruction and attention supervision fix different errors. FRM lowers the typical error; VisEQ removes the catastrophic ones.
Reconstruction at one pyramid level cuts validation MAE by 16.9%, and at all three levels by 28.5%. But test RMSE rises, from 85.40 to 108.63 and 91.45, so a few images still go badly wrong. Adding VisEQ brings test RMSE down to 38.68, 54.7% below the baseline, and gives the lowest test MAE.
Design variants and loss terms FSC-147-OCC, validation and test
Design variants
| Metric | No FRM | FRM · one level | FRM · all levels | FRM + VisEQ |
|---|---|---|---|---|
| Validation MAE | 15.83 | 13.16 | 11.32 | 11.60 |
| Validation RMSE | 54.38 | 54.51 | 48.12 | 35.40 |
| Test MAE | 14.42 | 13.77 | 11.90 | 11.42 |
| Test RMSE | 85.40 | 108.63 | 91.45 | 38.68 |
Reconstruction loss terms, added one at a time
| Metric | ℓ2 | + cosine | + Charbonnier | + VisEQ losses |
|---|---|---|---|---|
| Validation MAE | 13.88 | 12.18 | 11.32 | 11.60 |
| Validation RMSE | 78.67 | 48.88 | 48.12 | 35.40 |
| Test MAE | 13.24 | 12.38 | 11.90 | 11.42 |
| Test RMSE | 88.93 | 87.04 | 91.45 | 38.68 |
Lower is better. The cosine term brings the largest single gain in validation RMSE (78.67 to 48.88); the full objective is the most consistent across splits.
Reconstructed features land where the clean ones are
Figure 5. t-SNE of features at the three pyramid levels (256, 512, 1024 channels). Red: features of the occluded image. Green: features of the clean image. Blue: reconstructions. The occluded features sit apart from the clean ones at every level. At the finest level the reconstructions overlap them almost completely, and at coarser levels they stay close though more spread out.
Limitations
- Totals, not positions. FRM reconstructs features that give the right count, but the density inside a mask need not sit exactly where the hidden objects are.
- It needs the mask. The occlusion mask is an input. Without one, CountOCC behaves like an ordinary open-world counter and cannot estimate hidden instances. Predicting the mask jointly is future work.
- Synthetic occluders. The black rectangles are simple. The gains concentrate on occluded-region error while visible-region error is unchanged, which suggests the model is not exploiting the occluder's appearance, and the CrowdHuman results use natural occlusion.
- A small cost on clean images. On unoccluded FSC-147, CountOCC stays competitive but trails CountGD (test MAE 7.02 against 5.74): a trade-off between robustness to occlusion and accuracy when everything is visible.
Citation
If this work is useful to you, please cite it as:
@article{arib2025counting,
title = {Counting Through Occlusion: Framework for Open World Amodal Counting},
author = {Arib, Safaeid Hossain and Akter, Rabeya and Chowdhury, Abdul Monaf and
Sourov, Md Jubair Ahmed and Hasan, Md Mehedi},
journal = {arXiv preprint arXiv:2511.12702},
year = {2025}
}