Computer Vision in Archaeology

An introduction

Petr Pajdla

Institute of Archaeology, Czech Academy of Sciences, Brno

14. 9. 2026

The plan

1 Welcome · who is in the room
2 What computer vision is
3 Discussion 1 — your data, your question groups 1 & 2 report
4 The family of tasks
5 Discussion 2 — which task, and what does it cost? groups 3 & 4 report
6 Two ways of learning: discriminative & generative
7 Live demo — zero-shot classification with CLIP
8 Discussion 3 — what may you claim? groups 5 & 6 report
9 Wrap-up

Six groups of five

  • Please form six groups of five, sitting with people you do not already know
  • These groups stay together for the whole session — and are a good default for the week
  • Each round, two groups report back, so every group speaks once
  • Nominate a note-taker now; the reporter can change each round

Use your own material

Where a question says “your data”, use a project you actually work on. If you have nothing to hand, think about one of the datasets we are going to work with this week.

Who is in the room?

What you told us

Almost everyone has met machine learning; almost nobody does it. You know the vocabulary of detection and classification.

Two numbers worth pausing on

50%

use an AI coding assistant regularly – and 19% rate their Python 4 or 5 out of 5.

More of you lean on an assistant daily than feel fluent in the language it writes. That is not a problem to apologise for; it is the working condition of this week, and it is why Monday afternoon is about coding with agents rather than syntax.

80%

have never used an image annotation tool.

Annotation is probably the single largest cost in every project we will look at, and the one place where your expertise – not the model’s – decides the outcome.

What you want from the week

Multiple answers allowed. 48% have no dataset of their own to bring.

Confidently apply computer vision algorithms to detect archaeological features and micro-topographical anomalies, drastically reducing the manual processing time. (Árpád)

Better understand the inner workings of a machine learning model, by learning to develop and train it myself. (Aaricia)

Even if I don’t have a large corpus, I would like to test AI methods on my dataset by the end of the week. (Hüseyin)

What is computer vision?

A working definition

Computer vision is getting a machine to produce, from an image, the kind of structured description a person would produce by looking at it.

Three things hide in that sentence:

  • an image — which for a computer is only an array of numbers
  • a description — a label, a box, a mask, a measurement, a sentence
  • “the kind a person would produce” — someone decided what counts as correct

That last one is where archaeology enters. We define what the machine aims at.

An image is an array of numbers

A terra sigillata sherd, a 36 by 36 pixel window of it, and an 8 by 8 grid of the red-channel values.

Sherd: AMČR M-202300130-N00020, CC BY-NC 4.0

A 4000 × 3000 photograph of a sherd is 12 million numbers per colour channel — 36 million in all. Nothing in that array says sherd, terra sigilata, or Roman period.

The semantic gap

The same sherd as a colour photograph and as a map of brightness gradients.

Sherd: AMČR M-202300130-N00020, CC BY-NC 4.0

A specialist sees

  • a sherd
  • Roman, terra sigilata from Rheinzabern
  • 2nd half of the 2nd century CE

The computer receives

  • a 4000 × 3000 × 3 array
  • brightness gradients, edges, textures
  • no notion of sherd

Bridging that gap with learned statistical regularities is the whole enterprise.

Why now?

An aerial photograph of vegetation marks in a field, an excavation trench being photographed from above, and a 1984 archive photograph of a burial in a circular pit.

M. Gojda, AMČR C-DL-201000159; N. Profantová, AMČR C-200913914A-DT-05281; AMČR M-FT-900144861 (1984); CC BY-NC 4.0

  • Excavation photography is born-digital and effectively free — thousands of frames per season
  • Drone and satellite coverage is routine, not exceptional
  • Digitisation programmes produce collections nobody has time to look through
  • Pre-trained models mean you no longer need a million labelled images to start

The bottleneck moved from acquiring images to looking at them.

It is a measurement, not an interpretation

  • A model gives you what and where — never why
  • It reproduces whatever regularities are in its training data, including the unwelcome ones
  • It is a fast, consistent, literal-minded assistant who has seen your training set and nothing else

Everything a model outputs is a number you still have to interpret. That interpretation is your job, and it stays your job.

Discussion 1

Your data, your question

8 minutes in your group. Groups 1 & 2 report back.

  1. Describe one visual dataset you work with, or would like to. How many images? Who made them? Are they public?
  2. What question would you ask of it that currently takes too long to answer by hand?
  3. If a machine handed you that answer tomorrow, how would you check whether it was wrong?

Note

Keep your notes. Questions 2 and 3 come back in the next two rounds.

The family of tasks

Four questions, four output shapes

A Roman bronze coin with the single label coin.

Classification

What is this?

One label per image

Twenty-four potsherds on graph paper, each inside a bounding box.

Detection

What, and where?

Boxes + labels

A rosary with its chain, loose beads and cross each covered by a coloured mask.

Segmentation

Which pixels?

Per-pixel mask

A vessel profile drawing with landmarks at the rim, neck, maximum diameter, base and base centre.

Keypoints

Where are the landmarks?

Coordinates

The choice is not technical. It follows from the question you asked — and it determines how much annotation you will have to pay for.

The same photograph, five questions

Task Asked of one coin photograph Annotation cost
Classification Is this an antoninianus? one label per image
Detection Where is the legend? Where is the portrait? a box per element
Segmentation Which pixels are corrosion? a traced outline per region
Keypoints Where are the die-axis markers? a few points per image
Retrieval Which coins in the corpus look like this one? none — but you need a corpus

Cost rises steeply down the table. Start with the cheapest task that answers your question.

The same four shapes, across archaeology

Shape Published example
Classification decorated ceramic typology from photographs (Pawlowicz and Downum 2021)
Detection potsherds in high-resolution drone imagery (Orengo and Garcia-Molsosa 2019); burial mounds in historical maps (Berganzo-Besga et al. 2023); archaeological objects in LiDAR (Verschoof-van Der Vaart and Lambers 2019); illustrated objects in scanned catalogues (Klein et al. 2025)
Segmentation hollow roads traced in LiDAR (Verschoof-van der Vaart and Landauer 2021); agrarian field systems from satellite imagery (Küçükdemirci 2025)
Keypoints tracing a knapper’s movements, the motor side of knapping skill (Pargeter et al. 2022)
Retrieval 3D matching and retrieval of objects across collections (Jiménez-Badillo et al. 2023)
Berganzo-Besga, Iban, Hector A. Orengo, Felipe Lumbreras, et al. 2023. “Curriculum Learning-Based Strategy for Low-Density Archaeological Mound Detection from Historical Maps in India and Pakistan.” Scientific Reports 13 (1): 11257. https://doi.org/10.1038/s41598-023-38190-x.
Jiménez-Badillo, Diego, Omar Mendoza-Montoya, and Salvador Ruiz-Correa. 2023. “Application of Computer Vision Techniques for 3D Matching and Retrieval of Archaeological Objects.” F1000Research 12: 182. https://doi.org/10.12688/f1000research.127095.1.
Klein, Kevin, Antoine Muller, Alyssa Wohde, et al. 2025. “An AI-Assisted Workflow for Object Detection and Data Collection from Archaeological Catalogues.” Journal of Archaeological Science 179: 106244. https://doi.org/10.1016/j.jas.2025.106244.
Küçükdemirci, Melda. 2025. “Tracing Past Agrarian Field Systems Through AI-Based Analysis of Satellite Imagery in Konya, Central Anatolia, Türkiye.” Archaeological Prospection 32 (4): 881–97. https://doi.org/10.1002/arp.2001.
Orengo, H. A., and A. Garcia-Molsosa. 2019. “A Brave New World for Archaeological Survey: Automated Machine Learning-Based Potsherd Detection Using High-Resolution Drone Imagery.” Journal of Archaeological Science 112: 105013. https://doi.org/10.1016/j.jas.2019.105013.
Pargeter, Justin, Cheng Liu, Megan Beney Kilgore, Aditi Majoe, and Dietrich Stout. 2022. “Testing the Effect of Learning Conditions and Individual Motor/Cognitive Differences on Knapping Skill Acquisition.” Journal of Archaeological Method and Theory, ahead of print. https://doi.org/10.1007/s10816-022-09592-4.
Pawlowicz, Leszek M., and Christian E. Downum. 2021. “Applications of Deep Learning to Decorated Ceramic Typology and Classification: A Case Study Using Tusayan White Ware from Northeast Arizona.” Journal of Archaeological Science 130: 105375. https://doi.org/10.1016/j.jas.2021.105375.
Verschoof-van Der Vaart, Wouter Baernd, and Karsten Lambers. 2019. “Learning to Look at LiDAR: The Use of R-CNN in the Automated Detection of Archaeological Objects in LiDAR Data from the Netherlands.” Journal of Computer Applications in Archaeology 2 (1): 31–40. https://doi.org/10.5334/jcaa.32.
Verschoof-van der Vaart, Wouter B., and Juergen Landauer. 2021. “Using CarcassonNet to Automatically Detect and Trace Hollow Roads in LiDAR Data from the Netherlands.” Journal of Cultural Heritage 47: 143–54. https://doi.org/10.1016/j.culher.2020.10.009.

The pipeline for the rest of the week

Six stages in a row: question, data and annotation, pre-processing, training, evaluation, interpretation. The first two are highlighted.

Stage What you decide When
Question what would count as an answer? today
Data & annotation classes, vocabulary, licences Tuesday
Pre-processing scale, colour, augmentation Wednesday
Training architecture, transfer learning Thursday
Evaluation which errors you can live with Thursday–Friday
Interpretation what you may actually claim Friday

Projects fail at stage one or two. Almost never at stage four.

Good and bad questions for a model

Tractable

  • Which of these 8 000 photographs contain a scale bar?
  • Where in this ortho-image are the circular anomalies?
  • Which sherds share this decorative motif?

Not (yet) tractable

  • Is this settlement Iron Age? — from a photograph alone
  • Was this tool used on hide or on bone? — without a reference collection
  • Which of these finds matters more?

The test: could a competent colleague answer it from the image alone, consistently, in two seconds? If not, a model will not either — it will simply be confident about it.

Discussion 2

Which task, and what does it cost?

8 minutes. Groups 3 & 4 report back.

Take the question your group settled on in round one.

  1. Which of the four shapes does it map onto — classification, detection, segmentation, keypoints?
  2. Write the class list. How many classes, and be specific. What happens to images that fit none or two?
  3. Would two specialists independently give the same image the same label? Where exactly would they disagree?
  4. Estimate the annotation bill: how many images × how long each. Say the number out loud.
  5. Is there a cheaper task shape that would still answer your question?

Two ways of learning

Discriminative vs. generative

Discriminative

Learns the boundary between classes.

Given this image, what is the label?

Answers questions about existing data.

Generative

Learns what the data look like.

What does an image of this class look like?

Produces new data.

A discriminative model can tell a Neolithic pot from an Iron Age one. A generative model can draw you a Neolithic pot. Only one of them has ever seen the pot you care about.

The same idea, drawn

Two clusters of points, Neolithic and Iron Age, separated by a straight blue line.

Discriminative — draw a line separating the classes. Anything far from the line is ignored.

The same two clusters with orange density contours around each, and a few new points sampled from them.

Generative — model where each class lives. You can then sample new points from it.

Discriminative — the working horse of the week

A sherd photograph goes into a model, which returns a label, a bounding box, or a mask.

Sherd: AMČR M-202300130-N00020, CC BY-NC 4.0

  • Artefact classification — a typological class for a photographed find
  • Feature detection on LiDAR and satellite imagery — mounds, enclosures, field systems, hillforts
  • Use-wear segmentation — the pixels showing a particular polish or striation
  • Triage — separating scale-bar-only frames from find photographs
  • Retrieval — “show me the twenty most similar sherds in the corpus”

Common thread: the model points at evidence that already exists.

Generative — a smaller part of the picture

Two damaged medieval frescoes, each next to a version in which a neural network has filled in the missing paint.

Merizzi et al. 2024, Heritage Science 12, 41, Figs. 12 & 17, doi:10.1186/s40494-023-01116-x, CC BY 4.0, modified

Restoration

Sharpening or completing degraded imagery, as above

Synthetic training data

Manufacturing examples for rare classes

Vision-language models

Describing an image, answering questions about it

Common thread: the model produces something that was not in the record.

Side by side

Discriminative Generative
Learns the boundary between classes the distribution of the data
In → out image → label / box / mask prompt or noise → image / text
Typical models CNNs, YOLO, ResNet, SAM diffusion, GANs, VAEs, VLMs
Needs labelled examples many examples; labels optional
Fails by confident wrong labels plausible fabrication
Evidential status a measurement to verify an illustration or a hypothesis

The asymmetry that matters

A discriminative error is visible — the box is in the wrong place. A generative error is invisible — a fabricated rim looks exactly as convincing as a real one.

Where the line blurs

  • Vision-language models generate text — but you use them to classify: “is there pottery in this photo?”
  • Segment Anything (SAM) is a foundation model prompted like a generative one, producing discriminative masks
  • Synthetic augmentation uses a generative model to feed a discriminative one

The distinction is not a rigid taxonomy of architectures. It is about what the model outputs, and what you are entitled to claim from it.

That is exactly what the next eight minutes are for.

Live demo

Zero-shot classification with CLIP

You give the model an image and a list of candidate descriptions in plain English. It tells you which description fits best. No training. No annotation.

labels = ["a photo of a potsherd",
          "a photo of a coin",
          "a photo of a stone tool",
          "a photo of a scale bar"]

image = load("find_0042.jpg")
probs = clip_zero_shot(image, labels)   # → [0.81, 0.05, 0.11, 0.03]

Try to run the notebook yourself…

What just happened?

CLIP output for two sherds. Terra sigillata with five sensible labels: 95% potsherd. A coarse sherd with the same labels: 51% potsherd, 48% stone tool. The coarse sherd offered only coin, stone tool and bicycle: stone tool at 100%.

Sherds: AMČR M-202300130-N00020; M. Španěl, AMČR M-202300087-N00486; CC BY-NC 4.0

  • The model never saw an archaeological training set — it learned from image–caption pairs off the internet
  • It is being used discriminatively (a label out) but was trained the way a generative model is
  • Change the wording of a label and the answer changes — the vocabulary is now a parameter
  • Give it only wrong options and it will still pick one, confidently

Important

Zero-shot is a superb way to explore a collection. It is a poor way to report a number.

Discussion 3

What may you claim?

6 minutes. Groups 5 & 6 report back.

A colleague reports: “Our model detected 340 previously unrecorded burial mounds in the survey area.”

  1. What do you need to know before you believe that number?
  2. What is the difference between 340 detections and 340 mounds?
  3. What if the model was trained on mounds from a different region entirely?
  4. The paper includes a generated reconstruction of one mound. What must the caption say?
  5. Now apply all of this to the CLIP result you just watched. Would you put it in a paper? With what caveat?

Reporting back

Groups 5 & 6, then everyone in one sentence:

  • The dataset and the question you settled on
  • The task shape and the annotation bill
  • The one thing that would make it fail

Five things to carry into the week

  1. An image is numbers; a model is a very literal assistant
  2. The task shape follows from the question — and it sets the annotation bill
  3. Discriminative models point at evidence; generative models manufacture plausibility
  4. Both are useful; only one of them is a measurement
  5. Projects fail at question and data, not at training

Where to read more

The MAIA COST Action maintains a shared, open Zotero library on AI in archaeology — zotero.org/groups/5991067

Start here

  • Bellat et al. 2025 — Machine learning applications in archaeological practices: a review
  • Gattiglia 2025 — Managing Artificial Intelligence in Archaeology: an overview
  • Bickler 2021 — Machine Learning Arrives in Archaeology
  • Tenzer et al. 2024 — Debating AI in Archaeology
  • Calder et al. 2022 — Use and Misuse of Machine Learning in Anthropology

Worked case studies

  • Pawlowicz & Downum 2021 — ceramic typology from photographs
  • Orengo & Garcia-Molsosa 2019 — potsherd detection on drone imagery
  • Verschoof-van der Vaart & Lambers 2019 — R-CNN on LiDAR
  • Fiorucci et al. 2022 — evaluation measures for detection on LiDAR
  • Deligio et al. 2024 — ClaReNet, classifying Celtic coinage
  • Núñez Jareño et al. 2021 — synthetic data for Roman fine ware