Sketch Map Analyser

A Researcher's Guide

How to measure what people remember about a place — and what the numbers mean.

Start here

When you ask someone to draw a map of a place they know, what they produce is not a bad map — it is evidence. A sketch map externalizes a cognitive map: what the person encoded, what they kept, and how they organized it. The problem has always been turning that evidence into numbers you can analyse.

Traditionally this meant hand-coding: two raters, a scoring sheet, an inter-rater reliability check, and weeks of work per study. The Sketch Map Analyser replaces that with a set of automatic, reproducible measurements. You give it a reference map of the real environment and a participant's sketch map, tell it which drawn features correspond to which real ones, and it returns a row of numbers per participant.

What you do not need: any programming, any command line, any GIS experience. You work in a web page: load maps, click features to match them up, press Analyse, download a spreadsheet.

The three questions

Everything the tool measures is an answer to one of three questions about a participant's drawing.

1

What did they remember?

Of everything that exists in the real environment, how much appears in the drawing at all? This is a recall measure — it says nothing about whether things were placed correctly.

Completeness
2

How much did they simplify?

People collapse a row of shops into one block, or straighten a curved street. The tool identifies these simplifications so they are not mistaken for memory errors.

Generalization
3

Did they get the layout right?

Given what they drew, is it arranged correctly? Measured two ways: whether relationships hold (is the park still north of the station?) and whether distances and angles are accurate.

Accuracy measures

Keeping these separate is the central design decision. A participant who draws only three landmarks but places them perfectly is not the same as one who draws twenty landmarks in a jumble — yet a single "accuracy" score would rate them similarly. Here, the first has low completeness and high configural accuracy; the second the reverse. Both patterns are theoretically meaningful, and you can analyse them independently.

Key concepts and vocabulary

A few terms recur throughout the software and this page. They come from geoinformatics rather than psychology, so they are worth defining once.

Base map (or reference map)

The ground truth — an accurate map of the real environment your participants are drawing from memory. Usually derived from OpenStreetMap.

Sketch map

The participant's drawing, traced into the software so that each drawn object becomes a separate, countable feature.

Feature

A single object on either map: one building, one street segment, one park. The unit everything is counted in.

Landmark

A point or area feature — a building, monument, square. Distinguished from streets, which are line features and are counted separately.

Alignment

Your judgement about which sketch feature corresponds to which real feature. You make these matches by hand; every measurement depends on them.

Generalization

A simplification the participant made — merging several buildings into one, collapsing a roundabout to a crossing, drawing an area as a line.

Generalized base map

The reference map rebuilt at the participant's level of detail. This — not the original — is what accuracy measures compare against.

Junction

A point where two or more streets meet. Junctions can be analysed instead of buildings, which is useful for route-based or street-dominated sketches.

Why the "generalized" base map matters

This is the single most important idea in the tool, and the one most easily missed.

Suppose the real environment contains three adjacent shops. Your participant draws one rectangle and labels it "the shops." Scored naively, that is two omissions and one misplaced building.

But that is not a memory failure — it is categorical abstraction, and arguably a sign of well-organized spatial knowledge. So before any accuracy measure runs, the software merges those three shops on the reference map into one, matching what the participant did. The comparison is then like-for-like.

The consequence for interpretation: accuracy scores tell you how well someone laid out the world as they conceived it, not how well they reproduced a cartographer's map. If you want the latter, completeness is where the omissions show up.

Designing a study

What you need before you start

You needNotes
A defined study areaBounded and consistent across participants. Everyone should be drawing the same environment.
A reference map of that areaLoaded into the software from OpenStreetMap data.
Participants' sketch mapsDrawn on paper and traced in, or drawn directly in the editor.
Time for alignmentThe manual step. Budget realistically — see below.

Practical guidance

Standardize the drawing task. Sketch maps are highly sensitive to instructions. "Draw the campus" and "draw your route from the station to the library" elicit different representations (survey versus route knowledge). Whatever you choose, keep it identical across participants and report it.

Keep paper size and orientation constant. The tool normalizes for overall scale and rotation, but a participant who runs out of paper compresses one side of their map — which becomes a real distortion in the data rather than an artefact you can remove.

Expect alignment to dominate your time. Analysis itself takes seconds; deciding that "this blob is the library" takes judgement. For ambiguous drawings, consider having two researchers align independently and reporting agreement — the automated measures are perfectly reliable, but the alignment they rest on is a human coding decision like any other.

Decide in advance whether you are analysing buildings or junctions. Landmark-based measures need enough matched buildings to be stable; junction-based measures suit street-heavy or route sketches. You can run both, but choose your primary measure before looking at the data.

Sample size within a map. The configural measures compare every pair of features. With only two or three matched landmarks there are very few pairs, and the scores become unstable — a single misplaced building can swing them dramatically. Treat results from sparse sketches with caution, and consider reporting the number of matched features alongside each score.

Running an analysis

  1. Load your project — the reference map and one or more sketch maps.
  2. Tidy the geometry (optional). Hand-drawn streets often don't quite meet, or one street gets drawn as several strokes. The software can propose fixes — joining near-miss endpoints, merging broken segments — and you approve each one. Nothing is changed without your say-so. Worth doing if you plan to use junction-based measures, since those depend on streets actually connecting.
  3. Align the features — click a sketch feature, then the real feature it corresponds to.
  4. Press Analyse. Choose which measures to run:
    • Completeness — always included
    • Accuracy — the relational measures
    • Buildings GMDA — configural accuracy from landmarks
    • Junctions GMDA — configural accuracy from street junctions
  5. Read the results table — one row per sketch map, grouped by measure family.
  6. Download Results — a zip of spreadsheets ready for your statistics software.
Running more measures costs almost nothing in time, and you cannot go back and add a measure to an old export without re-running. Unless you have a preregistered reason not to, tick everything — then analyse only what you planned to.

Reading your results

Each measure below lists its range, which direction is "better," what it is actually sensitive to, and what to watch out for. Use the tabs to move between measure families.

How much of the environment made it into the drawing. A pure recall measure.

Landmark completeness

0–100% · higher = more recalled

The percentage of reference landmarks that appear in the sketch.

Interpretation. A memory-quantity measure, closely analogous to free-recall scores. Sensitive to encoding, exposure duration and familiarity.
Watch out: counted against the generalized reference map, so merging several buildings into one does not count as omissions. This is deliberate, but it means completeness is slightly more forgiving than a raw feature count would be.

Street completeness

0–100% · higher = more recalled

The same calculation for street segments rather than landmarks.

Interpretation. Often dissociates from landmark completeness. Route learners commonly recall the streets they travelled while omitting landmarks off-route; survey learners often show the opposite. Reporting both separately is usually more informative than the average.

Overall completeness

0–100% · higher = more recalled

The mean of the two measures above.

Watch out: an unweighted average. If your environment has 12 landmarks and 90 street segments, both still contribute equally. Check the two components before relying on the composite.

Whether the relationships between features are preserved, independent of metric precision. A sketch can be geometrically hopeless yet relationally perfect: the church still inside the square, the shop still on the left of the road, the streets still crossing in the same order.

The software extracts every such relationship from both maps and compares them one by one, across six families: containment and overlap, street–area topology, left/right of a road, relative direction between streets, street connectivity, and order along a route. Each family gets its own score; two summary measures aggregate across all of them.

Precision

0–1 · higher = better

Of all the relationships the participant expressed in their drawing, the proportion that match reality.

Interpretation. "When they committed to a spatial relationship, how often were they right?" Insensitive to what they left out — a sparse but correct map scores high.

Recall

0–1 · higher = better

Of all the relationships present in the real environment, the proportion the participant captured correctly.

Interpretation. Penalized by omission. Typically much lower than precision, because the reference environment contains far more relationships than anyone draws. Do not be alarmed by low values — compare across conditions rather than against an absolute standard.
Watch out: these are the standard information-retrieval definitions, not "recall" in the memory sense. Consider renaming them in your write-up to avoid confusing readers.

Per-family correctness

0–100% · higher = better

The same match rate calculated separately for each relationship family.

Interpretation. Useful for theoretically targeted predictions — for example, that route learners preserve sequential order along a path better than global direction. Look at the individual families rather than only the aggregate.
Watch out: families with very few relationships in a given map produce noisy percentages. The counts are exported alongside, so check them.

The metric measures, from the Gardony Map Drawing Analyzer — an established, published method for quantifying sketch map accuracy. It works by taking every pair of features and asking how the drawn relationship between them differs from the true one, then aggregating those pairwise errors.

Available in two variants: Buildings (landmarks) and Junctions (street intersections). Both produce the same six scores.

CanOrg — canonical organization

0–1 · higher = better

Whether features are in the correct broad direction from each other (north/south, east/west), calculated across all pairs in the real environment.

Interpretation. Because the denominator includes pairs the participant never drew, omissions lower this score. It is a combined measure of recall and coarse layout — the closest thing here to a single "overall quality" number.
Watch out: values are often strikingly low (0.1 is common) simply because participants draw a fraction of the environment. This is expected. It is a relative measure — compare conditions, not to 1.0.

CanAcc — canonical accuracy

0–1 · higher = better

The same directional check, but calculated only over the features the participant actually drew.

Interpretation. Layout quality with recall removed. Read alongside CanOrg: a large gap between them means the participant drew little but organized it well. Completeness already measures the recall side, so CanAcc is usually the cleaner dependent variable.

DistAcc — distance accuracy

≤1 · higher = better

How well relative distances between features are preserved, after equalizing for the overall size difference between drawing and reality.

Interpretation. Because it is scale-normalized, drawing a small map does not penalize you — only getting the proportions wrong does. Sensitive to local distortions, such as expanding a familiar neighbourhood relative to its surroundings.

ScaBias — scaling bias

signed · 0 = unbiased

The same distance comparison, but keeping the sign.

Interpretation. Positive means the participant systematically expanded distances; negative means they compressed them. Where DistAcc measures error magnitude, this measures direction — the signature of systematic distortion rather than noise.
Watch out: near-zero can mean either no bias or expansions and compressions cancelling out. Always report it with DistAcc.

AngAcc — angular accuracy

0–1 · higher = better

How well the angles between pairs of features are preserved.

Interpretation. A finer-grained companion to CanAcc. Canonical direction only asks "roughly north?"; this asks "at the right bearing?" A participant can score well on one and poorly on the other, which is itself informative.

RotBias — rotational bias

−180° to +180° · 0 = aligned

Whether the whole drawing is systematically rotated relative to reality. Positive is clockwise, negative counterclockwise.

Interpretation. One of the most theoretically loaded measures. Systematic rotation often reflects an alignment to a reference frame other than north — a dominant street grid, a habitual approach direction, or the orientation the environment was learned from.
Watch out: it is circular data. Do not average rotation across participants with an ordinary mean — +170° and −170° average to 0°, implying no bias when both were rotated nearly half a turn. Use circular statistics.

Bi-dimensional regression is a long-established method for comparing two spatial configurations. Conceptually it is ordinary regression generalized to two dimensions: it finds the single best transformation — one uniform stretch, one rotation, one shift — that maps the real configuration onto the drawn one, then asks how much of the drawing that transformation explains.

r — bidimensional correlation

0–1 · higher = better

How well the configurations correspond once the best transformation is applied.

Interpretation. Read like a correlation coefficient. High values mean the drawing is essentially the real layout scaled, turned and shifted; lower values mean the internal arrangement itself differs.

DI — distortion index

0 upwards · lower = better

The residual distortion, computed directly from r.

Interpretation. Zero is a perfect match. Reports the same information as r on a scale where bigger means worse — report one or the other, not both as independent findings.

phi — scale factor

around 1 = same scale

How much larger or smaller the drawing is than reality, overall.

Interpretation. Usually a nuisance parameter driven by paper size, but meaningful if drawing scale is manipulated or free.

theta — rotation

degrees

The overall rotation of the drawing relative to the reference.

Interpretation. Conceptually parallel to RotBias, estimated by a different route. The two can be reported as converging evidence — and again, they are circular data.

alpha1, alpha2 — translation

horizontal and vertical shift

How far the fitted configuration is displaced.

Interpretation. Almost always a nuisance parameter — it mostly reflects where on the page someone began drawing. Rarely reported.
A general caution on direction. Most measures here are "higher is better," but three are not: DI (lower is better), and ScaBias and RotBias, which are signed biases where zero — not the maximum — is unbiased. Check the direction of every measure before interpreting a correlation table.

A worked example

Suppose one participant's row comes back like this:

MeasureValue
Landmark completeness35.7%
Street completeness40.9%
CanOrg0.10
CanAcc0.89
DistAcc0.94
ScaBias−0.0001
AngAcc0.79
RotBias−30.2°

How to read it.

They drew roughly a third of the environment. That is unremarkable — most people do.

The gap between CanOrg (0.10) and CanAcc (0.89) is the story. CanOrg is low almost entirely because two-thirds of the environment is missing from the denominator, not because the layout is poor. CanAcc, restricted to what they actually drew, is high: what they remembered, they organized correctly. This is a selective but well-structured representation, and it would be a serious misreading to describe this participant as spatially inaccurate on the strength of CanOrg alone.

Distances are well preserved (0.94), and scaling bias is essentially zero — no systematic expansion or compression.

The interesting result is RotBias at −30°. Angular accuracy is noticeably lower than distance accuracy (0.79 against 0.94), and this rotation is why: the whole configuration is turned roughly 30° counterclockwise. That is not random error. It suggests the participant encoded the environment relative to some frame other than north — a main street's orientation, or the direction they habitually approach from. Worth checking whether it replicates across participants; a consistent rotation across a sample is a finding, not noise.

The general habit: never read a single score in isolation. Completeness contextualizes CanOrg; RotBias explains gaps between AngAcc and DistAcc; ScaBias qualifies DistAcc. The measures were designed to be read as a profile.

Getting your data into statistics software

Download Results gives you a zip of CSV files. CSV opens directly in Excel, SPSS, R, JASP and jamovi.

FileWhat it containsUse it for
ResultSummary.csvOne row per sketch map, every measure as a columnYour main analysis file
CompletenessDetailedOutput.csvThe underlying feature countsChecking what drove a completeness score
GMDADetailedOutput.csvConfigural measures plus how many features were matchedScreening for sparse, unstable cases
BDRDetailedOutput.csvRegression parameters per mapDistortion analyses
QADetailedOutput.csvRelationship counts per familyFamily-specific hypotheses
GeneralizationDetailedOutput.csvWhich features each participant merged or simplifiedStudying abstraction itself
QualitativeRelations/The full relationship list for each mapCustom or exploratory relational coding

ResultSummary.csv is already in the shape most analyses want: participants in rows, measures in columns. Add your condition and demographic variables as extra columns and you are ready to analyse.

Before you analyse. Sketch map measures are frequently non-normal — bounded scores pile up near their limits, and completeness is often right-skewed. Inspect distributions before defaulting to parametric tests, and remember that rotation measures require circular statistics.

Limitations and things to be careful about

The measures are computed reliably; that does not make every interpretation of them safe.

Alignment is a human judgement

Everything downstream rests on your decisions about which drawn feature is which. Reliable in clear cases, genuinely ambiguous in others. Treat it as coding, with the same safeguards.

Sparse sketches give unstable scores

Configural measures use feature pairs. Few matched features means few pairs, and single errors dominate. Consider a minimum-features criterion, set in advance.

Drawing ability is a confound

The tool measures the drawing, not the mental representation directly. Motor skill, confidence and willingness to commit to detail all contribute.

Scores are relative, not absolute

There is no established norm for a "good" CanOrg. These measures are for comparing conditions, groups or time points — not for classifying an individual.

Rotation and bias measures are circular or signed

Ordinary means and standard deviations mislead on RotBias and theta. Signed biases can cancel across participants, hiding real distortion.

Composite measures can obscure

Overall completeness averages landmarks and streets equally regardless of how many of each exist. Check components before interpreting composites.

Generalization handling is a choice

Applying participants' simplifications to the reference map is defensible and deliberate — but it does mean your accuracy scores are not directly comparable to methods that skip this step.

The environment shapes the numbers

A dense city centre and a sparse suburb yield different baseline values for the same participant ability. Do not pool across environments without care.

Research background

The tool was developed at the Spatial Intelligence Lab, Institute for Geoinformatics, University of Münster, building on a line of work concerned with making sketch map analysis systematic rather than manual.

Qualitative foundations

Establishing that sketch and metric maps can be aligned through qualitative spatial relations rather than coordinates — the basis of the relational measures (Schwering et al., 2014).

Classifying sketch maps

A feature-based scheme for distinguishing types of sketch map, informing how drawings are characterized before analysis (Krukar et al., 2018).

Understanding generalization

A systematic classification of how people simplify space when drawing (Manivannan et al., 2022), followed by a method for detecting those simplifications automatically (Manivannan et al., 2024).

Quantitative accuracy measures

Integration of the Gardony Map Drawing Analyzer and bi-dimensional regression, giving metric configural accuracy alongside the relational measures.

Key references

Technical reference

You do not need any of this to use the tool. It is here for research software engineers, for anyone installing it locally, and for readers who want to verify exactly how a measure is computed.

How the software is put together

Each analysis method runs as an independent service with its own address, coordinated by the main web application. This means a method can be updated or added without disturbing the others.

ComponentAddressRole
sketchmap_analyser:8000Editor, alignment, results table, export
generalizations:8001Builds the generalized reference map
completeness:8002Recall measures
qualitativerelations:8003Relational measures
validation:8004Geometry cleanup (researcher-approved)
gmda:8005Configural accuracy
bdr:8006Bi-dimensional regression

The main application calls these services from the browser — directly by port locally, and through a reverse proxy in production.

Installing and running it yourself

git clone https://github.com/ifgi-sil/SketchMapia-Microservices.git
cd SketchMapia-Microservices
docker-compose up --build

Then open http://localhost:8000/generalizingmaps/.

Production deployment uses prebuilt container images published automatically on each release, updated on the server by Watchtower and served behind an Apache reverse proxy that handles HTTPS.

Programmatic access (API)

Each service accepts map data as GeoJSON via HTTP POST and returns measures as JSON. Useful for batch processing outside the interface.

EndpointReturns
/generalizations/requestFME/The generalized reference map
/completeness/analyzeCompleteness/Completeness measures
/accuracy/analyzeQualitative/Relational measures and full relation lists
/validation/validate/Proposed or applied geometry corrections
/gmda/calculateGMDA/Configural measures (landmarks)
/gmda/calculateJunctionGMDA/Configural measures (junctions)
/bdr/calculateLandmarksBDR/Regression parameters (landmarks)
/bdr/calculateJunctionsBDR/Regression parameters (junctions)
{
  "CanOrg": 0.0962,
  "CanAcc": 0.8917,
  "ScaBias": -0.0001,
  "DistAcc": 0.9358,
  "RotBias": -30.2334,
  "AngAcc": 0.7942,
  "nTL": 14,
  "nDL": 5
}

nTL is the number of reference features; nDL the number drawn.

Exactly how each measure is computed

Full formulas, derivations and implementation notes live with the source code:

Measure familyDocumentation
Completenesscompleteness
Relational accuracyaccuracy
Generalizationgeneralizations
Configural accuracygmda
Bi-dimensional regressionbdr
Geometry cleanupvalidation
Interface and exportsketchmap_analyser

Source code: github.com/ifgi-sil/SketchMapia-Microservices