About Words Over Time

About / design + research

Words Over Time is an independent research archive tracing selected English words across time, media, and institutions. Its contribution is a repeatable method for comparing corpus signals, lexical attestations, archival occurrences, and public records without treating them as one dataset. Each study preserves source limits, denominators, date precision, and missingness, then turns those distinctions into an inspectable visual account of changing forms, uses, and public meanings—not a universal history of English.

Research method

Research object and scope

Each entry is a selected-word case study rather than a dictionary entry or a universal history of English. A study defines the forms being researched, the sources allowed to answer its question, the comparisons those sources support, and the gaps that prevent a stronger conclusion.

Research protocol

01Question and scope

Specify the headword and variants, time and geographic scope, research question, and comparison to be tested.

02Evidence register

Record each source’s role, release or capture date, fields, coverage, rights, and missingness before analysis.

03Analytic specification

Declare inclusion rules, denominator, grouping, transformation, unit, and incomparable states before rendering.

04Reproducible derivation

Calculate from retained inputs through scripts or typed transforms, preserving missing and unavailable values.

05Claim review

Match the conclusion to the evidence and publish its unit, caveat, source, access date, and revision path.

Evidence and measurement

AForever / form policy

forever and for ever remain separate one-gram and two-gram series. A combined rate is published only when same-release raw totals provide a shared word-token denominator.

BHub / fixed inventory

Visibility means selected phrases at or above 0.002 appearances per million, divided by the same 39 proxies in each of six twenty-year periods. It does not estimate all uses of hub.

CDepression / layered claim

The retained core series peaks at 43.33 appearances per million in 1932. Separate economic-phrase and NBER series support context, not a claim that every occurrence is economic or that timing proves causation.

Native denominators remain attached, including the distinction between unigram and bigram series. Smoothing, indexing, ranking, aggregation, and normalisation are labelled as transformations. Joined and spaced forms remain separate unless a valid common denominator supports combination. Missing, unavailable, not searched, and incomparable are never silently converted to zero.

Claim boundary and review

Claims are limited to named-corpus visibility, source-bound attestation, disclosed semantic grouping, reproducible transformation, and bounded interpretation. They cannot establish a universal history, equate frequency with importance or causation, infer first use from a first corpus point, or treat missing evidence as absence.

Review keeps event, text, publication, capture, and revision dates separate and checks source status, form policy, overclaiming, accessibility, and rights before publication.

Design

Mobile visual position

The mobile edition has no single national or period style. Its working references mix contemporary interface graphics, editorial reporting, data art, and environmental colour. Because the original reference ledger does not fully preserve authorship, the six named examples below are critical parallels for reading the finished pages—not verified sources of direct influence.

Across routes, scroll sets sequence, motion exposes structure, and colour creates atmosphere without replacing labels or source boundaries.

Applied references

01Data

Yugo Nakamura / iida UI

KDDI describes one continuous information band built from icons, widgets, section bars, motion, and vertical scroll. It is a useful parallel for Data’s modular stream—not a source for its charts or palette.

KDDI / INFOBAR A01
02Artificial

Ryoji Ikeda / datamatics

Ikeda treats data as number fields, grids, repetition, and precise audiovisual sequence. This parallels Artificial’s data-as-signal method; its magenta and layout remain project-specific.

MOT / Ryoji Ikeda
03Artificial

Rhizomatiks / Multiplex

Rhizomatiks turns source code, project metadata, event data, and motion into visual systems. This parallels Artificial’s instrument-like interaction, not a copied installation.

Rhizomatiks / Multiplex
04Hub / Bay Area

Barbara Stauffacher Solomon / supergraphics

SFMOMA describes supergraphics as large forms responding to architecture. This parallels Hub’s oversized Helvetica and page-as-environment, translated into a phone-length report.

SFMOMA / Supergraphics
05Hub / Southern California

Light and Space

Light and Space centres light, transparency, reflectivity, and colour. This parallels Hub’s non-quantitative chromatic atmosphere; values remain in labels and axes.

Getty / Light and Space
06Hub / MIT

Muriel Cooper / Information Landscapes

Cooper’s work moved from print toward dynamic information interfaces. This parallels Hub’s continuous, scene-based reading, not its specific colour field.

MoMA / Information Landscapes
Source

Evidence stays attached to its origin.

Source classes answer different questions. Every public claim retains source identity, coverage, date precision, transformation history, missingness, and rights rather than treating the archive as one interchangeable dataset.

Google Books Ngram Viewer

Long-run printed-book visibility in named English corpora through 2022.

Boundary: Viewer/API values retain their source release, n-gram order, native denominator, and transformation record.

Project Gutenberg + Library of Congress

Selected public-domain passages, scanned pages, and historical newspaper occurrences.

Boundary: Each occurrence retains its work, edition or publication context, OCR condition, date precision, and item-level rights.

Lexical references

Attestation and sense-history checks through cited dictionaries and etymological references.

Boundary: Entries are citation targets rather than reproduced datasets; a reported attestation remains source-bound and revisable.

Legal, clinical, policy, and technical sources

Domain anchors from courts, statutes, agencies, standards bodies, controlled vocabularies, and research publications.

Boundary: These records support bounded institutional context, not a shared frequency scale or automatic semantic classification.

Contemporary and aggregate records

Modern context from Wikimedia projects and selected geographic, demographic, academic, and public-attention sources.

Boundary: Captures remain source- and date-specific; a modern inventory does not establish persistence or prevalence.

Citation and provenance

Cite this archive and upstream sources separately. A chart records retrieval, cleaning, transformation, grouping, and visual design; it does not replace the source citation.

Pan, Dai. “[Word page title].” Words Over Time, 2026, [page URL]. DOI: 10.5281/zenodo.20437678. Accessed [day month year].

Example

Pan, Dai. “Hub.” Words Over Time, 2026, https://wordsovertime.com/words/hub. DOI: 10.5281/zenodo.20437678. Accessed 27 May 2026.

The DOI identifies the project, not a separate dataset assigned to every route.

Project DOI ↗
License

Software openness does not erase authorship or upstream rights.

The repository makes software and research processes inspectable. Original work may be cited and studied non-commercially with attribution; commercial copying, republication, dataset extraction, or reproduction of the finished identity requires permission.

Read the MIT licence ↗

Project source code

Application code, styles, utilities, and pipeline implementation are released under MIT; the grant covers software implementation only.

Original research and design

Research writing, curated datasets, classifications, page compositions, visual identity, and Dai Pan / 潘岱 authorship marks remain © 2026 Dai Pan / 潘岱.

Upstream material

Corpus series, archival passages, dictionary references, public records, metadata, and open media retain their own source terms and attribution requirements.

Processed and curated records

Derived tables and project records document provenance and transformations; they do not relicense upstream material or grant commercial extraction rights over curated classifications.

Site privacy

The public site uses no accounts, cookies, or visitor tracking and stores no visitor personal data.

Contact

Creator and correspondence

Research, data, writing, and design by Dai Pan / 潘岱, a Chinese artist, designer, and design researcher working across visual art, printmaking, writing, image-text worlds, and poetic research.

Questions, corrections, and source suggestions are welcome.

Emaildpan53853@gmail.comEmailjarl555@qq.comPhone+86 15262753021

Project repository

GitHub / dpan538Words-Over-TimePublic repository / code / data pipeline