
Making the data a model can learn from
Institute of Archaeology, Czech Academy of Sciences, Brno
Institute of Archaeology, Czech Academy of Sciences, Brno
Institute of Archaeology, Czech Academy of Sciences, Brno
15. 9. 2026
| 1 | The artefacts dataset — where the photographs come from | |
| 2 | What “annotated” means · COCO, YOLO, Pascal VOC | |
| 3 | Annotation tools · live demo of CVAT with SAM 2 | |
| 4 | You annotate, in your own CVAT project | 16 minutes |
| 5 | What you had to decide | discussion |
| 6 | Three problems — imbalance · vocabularies · missing metadata | |
| 7 | ArchaeoTag — everyone plays a round | phones out |
| 8 | Discussion — your dataset, your label set | groups 1 & 2 report |
Half of this session is you working, not me talking.
80%
have never used an image annotation tool.
Digital Archive of the AMCR – Individual finds selected

9 762
photographs
10 749
annotations
112
artefact types used
This is a real working dataset, with all the untidiness that implies. We will open the actual file in a minute.

Rosary: AMČR M-202400071-N00083, CC BY-NC 4.0
Same photograph, same objects. The annotation is a choice, and it is the expensive one.
| You want | You annotate | Roughly, per image |
|---|---|---|
| what is in this photograph? | one label | seconds |
| how many, and where? | a box per object | tens of seconds |
| what is its exact outline / area? | a polygon per object | minutes |
| what condition is it in? | attributes per object | seconds — if you decided the fields in advance |
The order matters
Going from boxes to polygons after you have annotated 2 000 images means annotating 2 000 images again. Going the other way is free.
Yesterday: the task shape follows from the question. Today: and the question has a cost.
{
"info": { ... },
"licenses": [ ... ],
"images": [ {"id": 1, "file_name": "...", "width": 3978, "height": 2762}, ... ],
"categories": [ {"id": 1, "name": "adornment"}, {"id": 63, "name": "fibula"}, ... ],
"annotations": [ {"image_id": 1, "category_id": 63, "bbox": [...], "segmentation": [...]}, ... ]
}annotations point at images and categories by id, so nothing is repeatedTip
Common Objects in Context — a benchmark dataset from 2014 whose file layout outlived it and became the de-facto interchange format.
{
"categories": [
{"id": 58, "name": "cross"}, …
],
"images": [
{"id": 8987, "file_name": "…/M202400071N00083F01.jpeg",
"width": 3978, "height": 2762}, …
],
"annotations": [
{"id": 9872, "image_id": 8987, "category_id": 58,
"bbox": [837.8, 1760.0, 485.2, 445.9],
"area": 216350.7,
"segmentation": [[1301.3, 1762.7, 1279.0, 1760.0, …]],
"iscrowd": 0,
"attributes": {"recognizable": "true", "occluded": false}}, …
]
}image_id points into images, category_id into categories — the three lists are joined only by idbbox is x, y, width, height — corner, not centresegmentation is one flat list: x1 y1 x2 y2 …, here 47 verticescategory_id is just 58. The word cross lives only in categoriesattributes is not standard COCO — the tool added itCOCO · one JSON file for the whole dataset
amcr-pas_v0.json
bbox is the top-left corner plus width and height, in pixels. The polygon is kept. The image and the class name are found only by id.
Pascal VOC · one .xml per image
Annotations/M202400071N00083F01.xml
The box is two corners, in whole pixels. The class name is written out, and the file names its own image. The polygon is gone.
Important
Every conversion is lossy in a different direction — and 57 means nothing without the data.yaml it came with.
| Tool | Runs | Good for |
|---|---|---|
| CVAT — cvat.ai | a server, or cvat.ai | images and video, a team, SAM 2 built in · what we use today |
| Label Studio — labelstud.io | a server | any data type — text, audio, time series as well as images |
| VIA — VGG Image Annotator | one HTML file in your browser | a small set, one person, no install, works offline |
| makesense.ai — makesense.ai | in your browser | quick boxes and polygons, no account, images stay on your computer |
They all do the same thing: shapes, classes, attributes → a file. Pick by team size and data type, and check the export formats before you start.
Why CVAT today
Live — the steps you will do next
labels.json: the classes and the three questions
labels.jsonis the schema. If you can read it, you know exactly what the dataset will contain.
Note
The AMČR-PAS annotations you have been looking at were made in CVAT — that is where recognizable and occluded come from.

Rosary: AMČR M-202400071-N00083, CC BY-NC 4.0
One photograph. At 10 749 annotations across the collection, that is the scale of what somebody already did by hand.
Watch it fail
A corroded find on a background close to its own tone, and the mask leaks into the scale bar. Model-assisted is not model-decided — correcting a bad mask can cost more than drawing a good one.
or this one, it is misbehaving:
Type it into your browser. It is a local address — it works only on the c network here.
About 5 minutes. Everyone works on the same 10 photographs, each in their own project.
http://10.10.1.40:8081 → Create an account2-tuesday-artefacts/labels.json · Done · Submit & Openimages/ under Select files · Submit & OpenTip
Stuck on step 3? The Raw tab must contain only labels.json — one [ at the start, one ] at the end.
16 minutes. Everything after this slide is about decisions you make now.
fragment a piece of an object, not a whole one?
damage none, minor or major?
scale is there a scale bar in the frame?
Tip
Do not compare with your neighbour. We will talk about what you decided, and why.
hands up Photo 04: coin or pendant? · Photo 06: mount or pendant? · Photo 08: one object, or several?
SAM made your outlines agree. It could not make your meanings agree — and that is the disagreement inside every archaeological dataset you will ever download.
Important
An annotation guide with worked examples of the hard cases is worth more than doubling your training set.

Fibulae and coins are 40.1% of every annotation in the dataset.

A classifier that ignores the image and always answers fibula scores 22%. Always fibula or coin: 40.1%.
Important
40.1% fibulae and coins is a fact about detectorists, reporting and archives. It is not a fact about what was in the ground.
A model trained here learns the collection process at least as well as it learns the artefacts.

The British PAS (FISH vocabulary) against AMČR-PAS: 841 terms, and four different ways to land on coin. Strap end is mordant — no string match finds that.
| means | safe to do | PAS → AMČR-PAS | |
|---|---|---|---|
skos:exactMatch |
interchangeable in any context | merge the two classes | 119 |
skos:closeMatch |
interchangeable in some contexts | merge, and say so in the paper | 53 |
skos:broadMatch / narrowMatch |
one is wider than the other | roll the narrow one up; never down | 232 / 7 |
skos:relatedMatch |
associated, not substitutable | do not merge | 17 |
| (none) | no counterpart | leave out, and count what you lost | 25 |
closeMatch does not chain
A closeMatch B and B closeMatch C says nothing about A and C. Two hops through a crosswalk and you have a class that means whatever the path happened to be.
SKOS is a W3C standard for publishing exactly this: concepts with stable URIs, their labels in several languages, and typed links between vocabularies.
@prefix skos: <http://www.w3.org/2004/02/skos/core#> .
@prefix fish: <http://purl.org/heritagedata/schemes/mda_obj/concepts/> .
@prefix amcr: <https://example.org/amcr-pas/> . # placeholder, see below
fish:96665 skos:prefLabel "Brooch"@en ;
skos:exactMatch amcr:fibula .
fish:140465 skos:prefLabel "Strap end"@en ;
skos:exactMatch amcr:mordant .
fish:97525 skos:prefLabel "Coin hoard"@en ;
skos:narrowMatch amcr:coin .
fish:95426 skos:prefLabel "Jetton"@en ;
skos:relatedMatch amcr:coin .fish: URIs are the FISH thesaurus’ own; amcr: is a placeholder — the sheet has AMČR-PAS labels, not URIsThe realistic outcome
You will lose resolution. Of the 112 AMČR-PAS categories in the dataset, only 63 have a validated exactMatch in PAS — and 28 PAS types roll up into instrument alone. That is still a far better dataset than either alone — and the mapping is a research output in its own right.
What the file records:
Phones out. Open the URL on the board — three minutes of play.
Tip
The guessing is the incentive. The metadata is the point, and nobody playing has to care.
The same logic as the last hour: disagreement is data, not failure.
Warning
Crowd labels are good for is there a scale bar in this photograph. They are not good for is this an Almgren 236 fibula. Know which question you are asking.
8 minutes in your group. Groups 1 & 2 report back.
Note
Question 2 in hours, out loud. It is usually the moment a project’s scope becomes visible.
Groups 1 & 2, then anyone with a good answer to (3):
Tip
This afternoon is yours: present your own data and problems. Everything on this slide is a good place to start from.
This afternoon — your own projects and datasets. Tomorrow — microscopy, and what to do to images before a model ever sees them.
Tools
Data and vocabularies
Computer Vision in Archaeology · Brno 2026