Methodology and Sources
Methodology and Sources for Itinavia.
Last updated: 2026-10-09
Every itinerary is produced by one pipeline. The order matters, because each stage only ever adds facts it can source.
1. Place selection. Places are pulled from OpenStreetMap within a radius of the city centre, restricted to categories people actually visit — museums, temples and shrines, castles, parks, viewpoints, markets, notable restaurants — and to objects that carry a Wikidata entry. That single filter removes most of the noise: a plaque, a bus shelter and a bandstand all have coordinates; only the ones with an independent existence elsewhere on the web are kept.
2. Ranking. Candidates are ranked by a notability score built from the type of place, whether it has a Wikimedia Commons category, and — the strongest signal — how many language editions of Wikipedia have an article about it. A place documented in eighty languages is not the same as one documented in three, and the plan should know that. Distance from the centre applies a penalty so that a museum in a neighbouring town does not outrank one in the city.
3. Clustering into days. Days are separated by geography before anything is chosen. Each day gets one district; a later day's stops are kept away from earlier days' stops. This is what stops the classic failure where a "three-day" plan sends you back to the same neighbourhood on the first and third mornings.
4. Sequencing and timing. Within a day the order is chosen by evaluating every possible order of that day's stops and keeping the shortest total route that respects time-of-day constraints. Time on site comes from a published-by-type dwell table — a museum is not a bus stop — and transfers are estimated from real distance.
5. Budget. Costs come from a single published reference table keyed by country, split into stay, food, admission and local transport, with admission charged per person. The five figures always sum exactly to the total you asked for. There is no flight figure, because the request never contained a departure city.
6. Copy. A language model is given the finished plan and asked only to polish wording. Anything it produces that asserts a price, a time, an opening hour or a link not present in the plan is rejected wholesale, and the deterministic text is used instead.
7. Verification. The finished itineraries are scored by a separate quality gate that checks six things on every city: no day crosses town beyond a sane radius, no district repeats across days, no time slot holds two places, no day exceeds a working day's length, no invented flight cost, and every place resolvable back to a real coordinate. The current twenty-city score is published in the project changelog.