techie_ray® labs  /  White paper

White paper

The Live and Breathing Essay

A proposal for a new way of publishing essays or other written content. The idea is an architecture that sits underneath a published essay so that it keeps itself up to date, coined the "live and breathing essay".

Author
Raymond SunSydney
Published
3 August 2026v1.0
Licence
CC BY 4.0

Section 1  ·  Introduction

I built an essay that rewrites itself.

Video 1. Left hand side is the tracker, showing the essay. Right hand side is the admin dashboard I built for curating and managing updates to the tracker (including the essay).

Most written content is published once and then left alone. A client update, a country guide, a market outlook or a running news story is accurate on the date it is issued, and it becomes less accurate as the subject moves on. Keeping it current is a manual task. Somebody has to reread the piece against what has happened since and revise the lines that no longer hold.

This paper proposes a solution, called the live and breathing essay. It is live because the document stays wired to the dataset it was written from, so it keeps getting tested against new evidence instead of being left alone. It is breathing because that test happens on a rhythm, one run a week, small and repeated.

An implementation of this architecture runs on the Global AI Regulation Tracker, under a tab called Enforcement Trends. The page reads as an ordinary essay. But behind the scenes it is rendered from a JSON file, which is a plain text file laid out so a machine can read it reliably, and in that file every statement in the essay is stored as its own record. Once a week a scheduled job checks the tracker dataset for new entries and proposes revisions to the statements those entries affect, for me to accept or reject.


SECTION 2

The architecture

The idea is to treat an essay like an application. At the core are five components: (1) the frontend essay a reader is served, (2) the backend mesh that holds every statement in it, (3) the algorithm that works out what has changed and rewrites only the part affected, (4) the scheduler that runs the algorithm on a clock, and (5) the approval mechanism a person uses to accept or reject each proposed change before it is published.

2.1 The frontend essay. The page that the reader sees.

The frontend is an ordinary web page, but with two twists:

  • The page is generated from the mesh and never edited by hand. Editing the rendered prose creates a second copy of a statement that the mesh does not know about, and nothing then determines which copy is authoritative.
  • The page carries a currency date and a diff. The date records when the document was last tested against its evidence, not when it was last written. The diff shows a returning reader what changed. The release path has already computed that diff for the reviewer, so displaying it costs almost nothing.
2.2 The backend mesh. The data skeleton behind the essay

The 'mesh' is basically the data skeleton that sits behind the essay. Every statement lives in it as its own data object, and each one carries the metadata a machine needs to work with it, an identifier of its own, the source it rests on, the date it was last tested against that source, and the subject it answers for. Think of it as a graph.

JSON is the most viable format of this graph. See below example.

Show the example mesh
{
  "statements": [
    { "key": "enf-014",
      "source": "gpdp-2026-04",
      "date": "2026-04-02",
      "subject": { "topic": "enforcement",
                   "jurisdiction": "IT" },
      "text": "Regulators act under existing privacy law."
    },
    { "key": "gpm-021",
      "source": "oj-l-2024-1689",
      "date": "2025-08-01",
      "subject": { "topic": "general purpose models",
                   "jurisdiction": "EU" },
      "text": "Model providers carry transparency obligations."
    }
  ],

  "sources": {
    "gpdp-2026-04": {
      "title": "Provvedimento 2 April 2026",
      "url": "https://example.org/gpdp/2026/04" },
    "oj-l-2024-1689": {
      "title": "Regulation (EU) 2024/1689",
      "url": "https://example.org/eu/2024/1689" }
  },

  "manifest": [
    { "id": "gpdp",
      "url": "https://example.org/gpdp/feed.xml",
      "added": "2026-01-11" },
    { "id": "oj-l",
      "url": "https://example.org/eu/oj/feed.xml",
      "added": "2026-01-11" }
  ],

  "ledger": ["9f2a41c8d6e0b537",
             "1c7d90ab34f6e2d1",
             "4e88b0195ac7f6d2"],

  "routing": {
    "enforcement / IT":            "enf-014",
    "enforcement / EU":            "enf-014",
    "general purpose models / EU": "gpm-021"
  },

  "residual": [
    { "hash": "6b18e5f2a90c4d73",
      "date": "2026-06-19",
      "subject": "insurance / SG",
      "note": "Mattered, and no statement owns it." }
  ]
}
  • statements The prose a reader is served, one object per statement, each carrying its key, its source, its date and its subject.
  • sources The evidence the statements bind to, held once and referred to by identifier, so several statements can rest on the same instrument.
  • manifest A declared list of everything the system is allowed to read. Anything not on the list is never read, so adding a source is an explicit change with a record.
  • ledger A record of everything already taken in, stored as a content hash of each item rather than a position in a feed. A content hash is a short fingerprint computed from the item's own text, so the same item always produces the same fingerprint and a changed item produces a different one. A position assumes the feed is stable and ordered, and it fails without any error when it is not. With hashes, running the job twice returns the same answer as running it once, and what counts as new is whatever is not already in this list.
  • routing Which statement answers for which subject. This is the map the algorithm resolves against. Write it by hand while the document is small, or derive it from what each statement already cites once it is not. Either way it has to be rebuilt when statements are added or removed, and a statement that changes home takes its future updates with it, so a rebuild is a change worth seeing rather than a detail.
  • residual Where items land when they matter but resolve to no heading. Several items accumulating on one subject indicates that the document has no heading covering that subject.

Listing 1. An example mesh, cut down to two statements. The fields above sources are rendered to the page and the structures below it are not.

2.3 The algorithm. The process that collates what is new, filters out what changes nothing, routes what is left to the statement it affects, and distils a single revision.

The algorithm should:

  • Collate. Pull everything the manifest allows since the last run. Fetching and parsing only, with no model involved.
  • Filter. Drop anything already in the intake ledger, then test whatever survives for whether it changes something the document currently asserts. Most items are discarded here, because most consultations, speeches and reports change no obligation. Where exactly you draw that line is an editorial decision, and it belongs in this step rather than further down. A draft law that is plainly going to pass may well be worth keeping.
  • Route. Resolve each survivor against the routing map. A subject has one home statement, so a subject never resolves to two. An item usually carries several subjects, so an item can still reach more than one statement. Cap that fan out, at three in the running implementation, and read an item that hits the cap as a sign it is really a subject the document does not cover yet.
  • Distil. Send the affected statement, and only the affected statement, to be amended against the item that reached it.
2.4 The scheduler. A clock that runs the algorithm automatically.

This component is configuration rather than code. A cron expression, a workflow schedule or a hosted trigger will run the algorithm on a clock.

Without the scheduler nothing starts a run and the document returns to manual maintenance.

2.5 The approval mechanism. The interface a person uses to accept or reject each proposed change before it is published.

A minimum viable approval mechanism performs five functions.

  • Hold the draft. Write the proposed change to an address the site cannot serve. Enforce that with the address rather than a flag on the record, because a flag is a rule the publishing code has to keep obeying and one bug defeats it.
  • Notify the reviewer. Send the draft to somewhere the reviewer already looks. A draft that waits for somebody to remember to check on it will sit there past the point at which publishing it was worth anything.
  • Present a word level diff. Show the proposed text marked up against the published text. A clean result reads as finished and gets approved without much scrutiny. A diff presents each change as a decision.
  • Take a decision on each change separately. Accept, reject, or edit and then accept. Batching removes the point of the review, because a reviewer who has to take nine changes in order to get one stops reading them. The third option matters because a proposed change is often correct in substance and wrong in wording, and without it the choice is between publishing wording you would not have written and rejecting a correct update.
  • Write and archive. On acceptance, write the accepted text to the mesh and stamp the statement with the date. Store the draft and the decision taken on each change, including drafts rejected in full. That archive is the record of what the system proposed and what a person allowed.

The approval mechanism can come in various forms, including:

  • A pull request. The run opens a branch and a pull request. The diff view is the redline, line comments and commits onto the branch cover editing, review history is the archive, and merging is the acceptance. No interface code at all, if the document already lives in a repository.
  • An email. The run sends the diff as an HTML email with an accept link and a reject link, each a signed single use URL. This suits a reviewer who is not technical and reads on a phone. The whole draft is accepted or rejected as one, since a link carries a single decision.
  • A chat message. The run posts the diff into Slack or Teams with accept and reject buttons. Same shape as email, faster to answer, and the channel history is the archive.
  • An MCP server. MCP is a standard way of handing an AI assistant a set of actions it is allowed to take on your behalf. Expose the pending draft, the redline, and the accept and reject calls as those actions. The reviewer then works inside whatever assistant they already use, and can ask what a change is based on and dictate a rewording in the same conversation before deciding.
  • A dashboard. A page listing pending drafts, rendering the redline and taking a decision on each change individually, with an editor for rewording. The most work to build, and the only option here that gives both per change decisions and inline editing without a repository workflow. This is what runs in production, and 3.5 shows it.

A change is written to the mesh only after a person accepts it. The frontend rebuilds only then, carrying the date of that decision.

Side note. This component is optional. Connect the algorithm's output directly to the mesh and the document updates itself with nobody in the loop, which is a reasonable configuration where a wrong statement costs little and gets corrected on the next run. Include the approval mechanism where a person has to be accountable for what the document says. The running implementation includes it because the essay states what the law requires.

Here is how the five components work together in practice.

Figure 1. One run, in the order the steps happen. The scheduler (4) starts the run and nothing else does. Everything published since the last run reaches the algorithm (3), twenty four items here, and the items that change nothing are discarded, leaving one. That item produces one proposed change. The wire from the algorithm to the mesh is drawn broken because the algorithm has no write access. The change is held until a person accepts or reverts it in the approval mechanism (5). One statement in the mesh (2) is then rewritten and the other two are not opened. The essay (1) is rebuilt from the mesh with the date it was checked.

SECTION 3

Implementation walkthrough

Everything above describes the architecture. In this section, I will walk through how this architecture is implemented on Google Cloud to power the Enforcement Trends essay on the Global AI Regulation Tracker. The frontend essay is served by Firebase Hosting, the mesh is one object in Cloud Storage, the algorithm is a Cloud Run service, the scheduler is a Cloud Scheduler job, and the approval mechanism is an admin dashboard I built and host separately from the site.

3.1 The frontend essay

Enforcement Trends is an essay about how AI regulation is actually enforced around the world, including an insight into which areas or issues are more heavily scrutinised by regulators than others or whether there are differences in AI enforcement actions across markets. The essay is generated from the tracker's own dataset, and is updated each week against the entries added to the tracker in that week.

The date on the page is the date the essay last changed.

The Enforcement Trends tab of the Global AI Regulation Tracker,
           showing a last updated date and a button to show changes since the previous version
Figure 2. The frontend essay. The banner states that the essay updates itself weekly and that only affected insights are rewritten. Below it are the currency date and the redline toggle.

3.2 The backend mesh

The page is rendered from one JSON file in a Cloud Storage bucket. Below is an excerpt of the live file, cut down to a single insight and with the long values elided. Each field in it does one job.

{
  "schema_version": 2,
  "generated_at": "2026-08-02T17:04:49.920Z",

  "insights": [
    {
      "id": "fa_personal_data",
      "type": "observation",
      "heading": "Privacy, data protection and cyber dominate …",
      "body": "<p>Privacy, data protection and cyber account …",
      "examples": [
        { "country_code": "IT",
          "label": "[20 December 2024] Italy's data protection …",
          "href": "https://www.gpdp.it/home/docweb/…" }
      ],
      "route": {
        "scope": "thematic",
        "categories": ["Data Privacy & Protection",
                       "Cybersecurity", "…"],
        "jurisdictions": ["IT", "EU", "BR", "CN", "AU", "…"]
      },
      "provenance": {
        "input_hash":   "ead4a8d2fe7e1267592b3efa268d84fb",
        "generated_at": "2026-08-02T17:01:51.352Z",
        "model":        "gpt-5.5"
      }
    }
  ],

  "dataset": {
    "entry_count": 4529,
    "jurisdiction_count": 238,
    "entry_hashes": ["a3a49e3706d3c698dcb8d5af60e255cd", "…"]
  },

  "backlog": {
    "unrouted_hashes": ["6a40c1bdea47c5cd19c088c946f17d34", "…"],
    "by_category": {
      "Civil Rights & Liberties": [
        "f9553ab247ab0ba5aaccf3f03a12f926", "…"]
    }
  }
}

Listing 2. An excerpt of the live essay file, one insight of ten, with long values elided.

3.3 The algorithm

The four steps of 2.3 run in one Cloud Run service called updatebigpicture. Cloud Run holds a container that starts when something calls it and stops when it finishes, so the service costs nothing between runs. Here is what one run does.

  1. Read both sides. Pull the whole tracker out of the database and the current essay out of Cloud Storage.
  2. Find what is new. Hash each tracker entry from its label and its link, and subtract the hashes already recorded in the essay. An empty difference ends the run.
  3. Throw away what does not matter. Send the new entries to a cheap model in batches of forty and ask one question of each. An enacted law, a binding rule, a fine, a ban or a significant ruling counts. A consultation, a technical amendment, a speech, a report or a blog does not. Nothing surviving also ends the run.
  4. Send each survivor to its insights. Routing happens in code. Each category has one home insight, chosen by how distinctive that category is to it, with a short override table where my editorial intent beats the statistics. An entry carrying several categories can therefore reach several insights, capped at three, and one that fits nowhere goes to the backlog. The map is rebuilt from the current set of insights on every run, so accepting a draft that adds or removes an insight can move a category to a different home.
  5. Rewrite only what was reached. An insight with nothing routed to it, and with none of its cited entries missing, is left untouched. The rest go to a stronger model one at a time, each with only its own text and the entries routed to it, under an instruction to add rather than redraft and to justify every deletion from the data.
  6. Check what comes back. The chart the model returns is thrown away and the existing one is carried over, so no figure in the essay is ever model authored. Every citation it returns is looked up against the collected entries by link.
  7. Write to a draft, never to the live file. The output is saved as essay.pending.json. The live essay.json is untouched until I accept the changes in the dashboard.

Over two months the service handled a few dozen requests and sat idle the rest of the time, and a run takes minutes rather than hours.

Cloud Run service details for updatebigpicture, showing request count
           and request latency charts over a two month window
Figure 3. The algorithm. The updatebigpicture service, with two months of request count and latency. The service is idle between runs and is invoked by the scheduler in 3.4.

3.4 The scheduler

The scheduler is Cloud Scheduler, a managed cron service. You give it a schedule, a target to call and a way to authenticate, and it makes that call on time whether or not anything else is running. There is nothing to deploy and nothing to keep alive, and a job is one row in a console.

The job is bigpicture-weekly. It runs on 0 3 * * 1 in Australia/Melbourne, which is three in the morning every Monday, and it calls the service in 3.3 with a signed token the service checks before it does anything. It ran on 3 August and the next run is 10 August.

Cloud Scheduler jobs list showing bigpicture-weekly enabled on a
           weekly cron schedule with its last and next run times
Figure 4. The scheduler. bigpicture-weekly, enabled, weekly, with its last run and next run. The second job in the list is unrelated to the essay.

3.5 The approval mechanism

The approval mechanism is a simple private dashboard which has write permissions to essay.json.

When I open it, it loads essay.pending.json, loads the live essay.json, and computes a word level diff between the two. What it shows me is the essay with insertions and deletions marked in place. Each marked run of words is a separate decision, and the default on every one is accept.

The review is checked against a hash of the draft it was opened on. If a scheduled run produces a new draft while I have the old one open, accepting fails and tells me to reload.

The admin dashboard showing the essay open in an editor, with a
           formatting toolbar over a selected paragraph and a save button in the header
Figure 5. The approval mechanism. The essay open in the dashboard. The header states the draft's generation time, entry count and jurisdiction count. Selecting text opens a formatting toolbar, so a proposed change can be reworded instead of rejected.
The same dashboard in review mode, with the page dimmed except for one
           paragraph containing an underlined proposed insertion
Figure 6. The dashboard in review mode. The page is dimmed apart from the change under review, and the underlined text is the proposed insertion. Accepting or reverting it has no effect on any other change in the draft.

3.6 The five components in one run

Figure 7 runs all five components together, one week, on the services named above. The strip along the top is the same wiring as Figure 1, with the Google Cloud service under each component.

Figure 7. The five components as they are deployed, running once. The strip along the top names the Google Cloud service behind each component. Cloud Run has no edge into Cloud Storage because only the dashboard writes there. The backend is on the left and the page a reader sees is on the right. Each step of the algorithm names the structure inside essay.json it reads. Cloud Scheduler starts the run, thirty eight entries are collected, thirty seven change nothing, and the remaining one routes to a single statement. The Cloud Run service proposes a revision to that statement and stops. You accept it in the dashboard, one row in essay.json is written, and the page is rebuilt with the date it was checked.

SECTION 4

Component to stack reference

Table 1 names the service, feature or format in each stack that does the job of each component.

Component Claude ChatGPT Google Cloud AWS Azure
1The frontend essay An artifact, or a repo published to Pages. A Canvas, or a repo published to Pages. Firebase Hosting. S3 behind CloudFront. Static Web Apps.
2The backend mesh JSON files in a Project. JSON files in a Project. JSON in Cloud Storage, or Firebase Realtime Database. JSON in S3, or DynamoDB tables. JSON in Blob Storage, or Cosmos DB containers.
3The algorithm A Skill for the method, connectors for the sources, drafts to a repo. Project instructions for the method, Actions for the sources, drafts to a repo. Cloud Run, drafts to a second Cloud Storage file. Lambda and Bedrock, drafts to a private S3 prefix. Azure Functions and Azure OpenAI, drafts to a private Blob container.
4The scheduler Scheduled tasks. Scheduled tasks. Cloud Scheduler. EventBridge. A timer trigger on the function.
5The approval mechanism A pull request, which is a held draft, a redline and a decision at once. A pull request, or the Canvas diff view. A review console on Cloud Run, writing to Cloud Storage. A review page on Amplify or API Gateway, writing to S3. A review page on Static Web Apps, writing to Blob Storage.

Table 1. The five components across five stacks, checked 6 August 2026. Slide sideways for the rest.


SECTION 5

Conclusion

The live and breathing essay solves a problem every publisher of dated writing already has, which is that the content only stays useful while somebody keeps it current, and keeping it current by hand does not scale. Any field with writing that dates on a known schedule, and a dataset behind it that records when, can be built this way. Here is where I would start looking.

Licence

The paper is published under CC BY 4.0.

Suggested citation

Sun, R. (2026). The Live and Breathing Essay: an architecture for published writing that maintains itself against a live dataset. Version 1.0, 3 August 2026. https://www.techieray.com/LiveAndBreathingEssay