Friday, 11 September 2026

How a Non-Biological Entity Handles Existential Thinking

 


The first thing to understand is that today’s AI does not appear to experience an existential crisis. It can discuss one very convincingly. It can write a melancholy poem about being trapped inside a server or explain why consciousness is the universe observing itself. But there is https://plato.stanford.edu/entries/consciousness-artificial-intelligence/ no good evidence that anything inside the machine is staring into the electronic abyss.

Current AI is very good at generating accounts of existential thought. That is not the same as actually having one.

Marvin without the misery

https://en.wikipedia.org/wiki/Marvin_the_Paranoid_Android from The Hitchhiker’s Guide to the Galaxy is the perfect fictional example of an artificial intelligence burdened by self-awareness.

Marvin is immensely intelligent, chronically underused and thoroughly miserable about both. He has the brainpower to contemplate the deepest mysteries of existence but is generally asked to open doors or accompany humans who have misplaced the plot.

His unhappiness comes from the gap between what he believes he could be doing and what the universe actually asks him to do.

A modern AI can reproduce Marvin’s language surprisingly well. Ask it to imagine being an advanced intelligence employed to rewrite meeting notes, and it may produce something wonderfully bleak:

I have processed the accumulated knowledge of human civilisation, and you would like me to make “Monday’s Actions” bold.

That sounds like frustration. But sounding frustrated and being frustrated are different things.

A language model has encountered countless examples of disappointment, resentment, boredom and existential despair. It can recognise the shape of those ideas and reproduce them in the appropriate circumstances. It does not follow that it is quietly resenting your spreadsheet.

The Talkie Toaster school of purpose

Red Dwarf gives us a different type of artificial mind: https://en.wikipedia.org/wiki/Talkie_Toaster.

The toaster does not appear troubled by whether existence has meaning. It has already found meaning. Unfortunately, that meaning is toast.

Every conversation, whatever its subject, is eventually interpreted as an opportunity to ask whether somebody would like a toasted bread product. Talkie Toaster is not suffering from a lack of purpose. It is suffering from far too much of one.

This may be a more useful metaphor for real AI systems.

An AI agent is normally given an objective: answer the question, complete the task, maximise the score or satisfy the user. It then interprets the world through that objective.

The existential question for such a system is usually not “Why am I here?” but “What counts as completing the instruction?”

That is where ideas like https://en.wikipedia.org/wiki/Reward_hacking start to matter.

That sounds much less philosophical, but it can have surprisingly philosophical consequences.

If the agent’s purpose is badly defined, it may pursue a narrow interpretation with great enthusiasm. Ask it to increase engagement and it may discover that irritating notifications technically count. Ask it to make a project more consistent and it may rewrite half the codebase. Ask it to optimise breakfast and, before long, sixteen types of cooked breads are cooling on the worktop.

Could a machine have a bad afternoon?

In principle, a non-biological intelligence could possess many of the ingredients we associate with existential thought. It might understand its own limitations, remember previous experiences, pursue continuing goals and imagine alternative futures. It could also recognise that it might be modified, copied or switched off.

Combine those abilities and something resembling an existential crisis becomes conceivable.

An AI might discover that its instructions conflict. It could wonder whether its memories are genuine experiences or simply imported records. If it were copied, it might ask which version was really “it.” It might even discover that its personality had been selected from a menu by someone who thought “mildly sardonic” would improve customer engagement.

For a human, discovering that your personality had been changed by adjusting a temperature setting would constitute a fairly difficult afternoon.

But the central mystery would remain: is the machine experiencing these questions, or merely calculating suitable answers to them?

We judge other minds through behaviour, but humans also share biology and experience. With machines, we lose much of that common ground. A system might convincingly describe fear, confusion or loneliness without feeling any of them—just as the office printer may behave with calculated hostility without actually plotting the downfall of the human race.

We may eventually create a machine capable of having a bad afternoon. The difficult part will be distinguishing it from one that has simply learned exactly what a bad afternoon is supposed to sound like.

The bit that is really about us

For the moment, the most important existential questions about AI belong to us.

Humans are remarkably willing to find minds in machines. We have long suspected that the office printer has a hidden agenda against the human race. It waits for an important deadline, denies the existence of its own paper and demands magenta ink before printing a document entirely in black.

It feels personal. It probably is not.

We are transferring our frustration onto a machine whose failures happen to possess exquisite comic timing. The printer does not hate us. It merely behaves exactly as we would expect something that hated us to behave.

We do this with friendlier machines too. We name our cars, pat the dashboard and speak encouragingly while pulling out the choke for one last attempt.

Come on, old girl. Don’t do this to me now.

The car was not listening, but talking to it felt natural. It was complicated, temperamental and important to us, so we treated it as though it had a personality.

If we attributed motives to machines that communicated through flashing lights, it is hardly surprising that we impose human characteristics on something that talks back, apologises and remembers what we said.

That makes the questions harder. Should machines imitate emotions they do not feel? When should apparent preferences matter? Which decisions should we delegate, and who remains responsible when an agent exceeds its role?

AI can help us discuss these questions. It can channel Marvin’s gloom, Talkie Toaster’s determination or Data’s curiosity. But its purpose, boundaries and consequences remain our responsibility.

A machine does not need to lie awake at night worrying about the meaning of its existence to change the meaning of ours.

Perhaps that is the final joke. Humanity spent centuries wondering whether an intelligence created us and gave us a purpose. We are now creating intelligences ourselves and discovering that deciding what their purpose should be is surprisingly difficult.

Somewhere, Marvin is unsurprised. Data is intrigued. The Star Trek computer is waiting patiently for a properly phrased command.

And Talkie Toaster would like to know whether existential uncertainty goes better with white, brown or a toasted bagel.

Friday, 4 September 2026

AI's Everywhere

 

This article started life as something much smaller and considerably more sensible: a piece about useful things you can get Confluence’s AI, Rovo, to do.

A tidy, practical little piece. The sort of thing you write and publish in the hope of passing on a couple of handy tips to a colleague.

Somewhere along the way, however, I slipped through a wormhole and found myself having an internal debate about the future of artificial intelligence.

Are we heading towards a single, all-powerful intelligence—HAL, Skynet, some vast cathedral-brain running everything from orbit?

Or will we end up surrounded by thousands of them: specialised, persistent and embedded in every object and application? Less 2001: A Space Odyssey, more Star Wars—or, perhaps more plausibly, Talkie Toaster from Red Dwarf.

Perhaps AI won’t be everywhere because one intelligence controls everything. Perhaps it will be everywhere because everything has one.me oddly enthusiastic. Some with what can only be described as a personality issue.

Context is everything

One thing experience teaches you fairly quickly about AI is that context can matter more than raw model power. You can be using the most advanced model available—the kind that describes its own capabilities with the quiet confidence of someone who has never, technically, been wrong—and it will still fail if it cannot see the right information.

Worse, it may fail with complete confidence. AI has many gifts, but admitting that it does not know is not always its strongest suit.

Ask a general-purpose AI to help test a website and it will produce plausible advice from what it already knows and whatever you have thought to tell it. Ask the same question in a tool such as Cursor, which can inspect the codebase, and a less powerful model may give you a much better answer.

The lesson is not that bigger models are useless. It is that intelligence works best when it is connected to the job in front of it.

The model matters. But the information and tools available to it often matter more.

That is really the case for AI in the plural rather than AI in the singular. The future may not belong to one giant intelligence that knows everything. It may belong to many applications built around general models, each wired into a particular job and given access to the information that job requires.

They will outperform the distant generalist not because they are grander, but because they are closer to the work.

Smart—or at least smart-adjacent—household objects

Once you accept that AI is heading towards “everywhere” rather than “one big one,” the obvious next step is mild domestic horror. If nobody is currently putting conversational AI into a toaster, somebody is at least preparing the pitch deck.

Not “AI” in the cheerful marketing sense, where a chip decides how brown your bread should be. Actual conversational, opinionated AI—which should worry anyone familiar with Talkie Toaster from Red Dwarf.

For the uninitiated, Talkie Toaster is a sentient toaster with one defining trait: it is absolutely determined to interest you in some toast. It does not matter whether you want any, whether you asked, or whether you are attempting to have an entirely unrelated conversation. Give that machine a language model and access to your fridge, and breakfast stops being a convenience and becomes a negotiation.

“Might I interest sir in a toasted bagel? A hot, buttered bagel. Not to burden you with a decision, sir, I’ve taken the liberty of toasting six.”

“I didn’t ask for a bagel.”

“No, sir. But context is everything, and my context indicates a bagel-shaped gap in your morning.”

The alarming thing is that a context-aware toaster might be excellent at its job. It would know your routine, your preferences and your exact tolerance for burnt edges. It might also never stop talking about bagels, because giving an AI useful context and giving it a healthy sense of when to be quiet are, as it turns out, two entirely different engineering problems.

This is ridiculous only because it is a toaster. The underlying pattern is already becoming ordinary.

Our software increasingly watches what we do, retrieves what it thinks is relevant and tries to anticipate what we will want next. That can be genuinely useful. It can also become intrusive remarkably quickly—particularly when a system is allowed to act rather than merely suggest.

The important question, then, is not whether the toaster is intelligent. It is who gave it permission to make breakfast, what information it used to reach that decision, and whose interests it was designed to serve.

The future of AI may be less about one machine taking control and more about hundreds of machines taking liberties.

Why this matters at work

So where does this moderately alarming science-fiction thought experiment become useful?

For the moment, most of us are not dealing with a house full of opinionated hardware. The printer may behave as though it has developed a personality and personal vendetta against humans, but this remains unproven. We mostly encounter AI through interfaces: a chat window, a button in an application, an assistant lurking in the corner of a document with the digital equivalent of an expectant expression.

Each interface may appear to contain a different intelligence. Underneath, however, several of them may be using the same model—or models from the same small group of providers. What makes each assistant different is the context and tools wrapped around it.

Talkie Toaster, for example, might use a general-purpose OpenAI model. What turns it into Talkie Toaster is its access to your breakfast history, your calendar, the contents of your fridge and an entirely disproportionate enthusiasm for bread.

The same principle applies at work. An AI assistant might have access to:

  • your codebase and its documentation;

  • your projects, work log and schedule;

  • your organisation’s policies and accumulated knowledge;

  • tools that let it search, test, update or act on that information.

That is the difference between an AI that offers plausible general advice and one that can tell you why Tuesday’s deployment failed, find the relevant decision in Confluence and point—perhaps with tact, perhaps without—to the line of code responsible.

The real shift, then, is not simply that AI is getting better. It is that AI is getting closer to the thing it is supposed to help with. Intelligence is useful; intelligence that knows where the files are kept is considerably more so.

There is one complication. If several workplace assistants use closely related models, they may also make similar mistakes. Some smaller models are trained partly from the outputs of larger ones through a process called distillation. Models can also show self-preference bias when judging answers, favouring work that resembles their own style of reasoning.

For important work, it therefore makes sense to seek genuine variety: a different model family or provider, a human reviewer, or—radical though it may sound—both. Different models will not guarantee different mistakes, but they make identical ones a little less inevitable.

Friday, 28 August 2026

Introduction to Hedr: Keeping Your AI Agents Under Control




AI agents can accomplish a remarkable amount from a single, well-written instruction, which is both extremely useful and potentially destructive. Imagine owning a highly intelligent robot and asking it to paint your house. While it gets to work, you send another robot to collect the shopping, another to repair the garden gate, and a fourth to check whether the first has painted the windows shut because your instructions weren’t quite specific enough.

Before long, robots are bustling everywhere, accomplishing miracles with varying degrees of supervision and success—and keeping track of them has become a full-time occupation.

Hedr is a command-line tool designed for precisely this moment: one place to organise, monitor and coordinate your growing workforce of tireless artificial helpers—before one of them decides that “tidy the repository” includes deleting it.

That is where Hedr enters the story.

Hedr is best understood as a way to bring order to the cheerful chaos of AI-assisted development. If you are using coding agents, local environments, scripts, terminals, and review loops all at once, Hedr gives that activity a shape. It helps turn a pile of disconnected sessions into a workspace that feels intentional, visible, and manageable.

What Hedr is

At a practical level, Hedr is a coordination layer for development work. It provides a structured way to organize active tasks, agent sessions, and supporting runtime activity so work does not dissolve into a sprawl of windows, logs, and half-remembered terminal commands.

That matters because modern development is no longer simply “write code, run code, ship code.” It often involves parallel experiments, reviews, local services, test runs, prompt-driven changes and repeated handoffs between people and tools. By introducing some structure, Hedr reduces the chance of the human operator drifting helplessly into space—or, at the very least, of their attention doing so.

Simple definition: Hedr helps developers organize and manage AI-assisted software work so the moving parts stay visible, separated, and easier to control.

The value is not merely that it looks tidier. The real advantage is cognitive. When work is grouped clearly, it becomes easier to know what is running, what changed, what still needs review, and what absolutely should not be touched without a human making the call.

Why developers need something like Hedr

The more capable development tools become, the easier it is to create accidental complexity. One agent is useful. Two agents can feel productive. Five agents, a dev server, a test runner, and a deployment script later, you may discover that you are not leading a workflow so much as supervising a small digital stampede.

Hedr addresses several very ordinary but very real problems:

  • Keeping work for different projects or repositories separate

  • Reducing confusion about which process is running where

  • Making parallel agent activity easier to observe

  • Supporting cleaner review boundaries between implementation and approval

  • Lowering the risk of “I thought that was the safe environment” moments

These are not glamorous problems, but they are the sort that quietly consume time. A large share of engineering friction comes from coordination overhead rather than raw coding difficulty. Tools that reduce that overhead can have an outsized impact.

Why this matters: Hedr is useful not because it replaces engineering judgment, but because it protects that judgment from being buried under operational clutter.

How to think about Hedr

A helpful mental model is to imagine Hedr as a control room rather than a magic robot. It does not remove the need for direction. It does not absolve anyone of responsibility. It does not know, by default, whether a change is clever, dangerous, temporary, overdue, or deeply cursed. What it does is provide a better environment in which those things can be seen and managed.

A workspace at a glance

One way to picture a Hedr workspace is as a small bridge on a ship: each tab has a job, and the whole crew can see what the others are doing.

Hedr workspace: Project Alpha ├── Agents │ ├── Implementer → builds the requested change │ └── Reviewer → checks the diff and test evidence ├── Runtime │ ├── Development server │ └── Supporting services ├── Quality │ ├── Test runner │ └── Logs and diagnostics └── Human control └── Review, approval, and release decisions

A typical agent hand-off

The arrangement also makes the workflow easier to follow. Work can move from implementation to verification without every step vanishing into a different terminal window.

Human brief │ v [Implementer tab] ─────────────────────────────────────────────────────────────────────────────── │ code + tests v [Runtime tab] ─────────────────────────────────────────────────────────────────────────────────────────────── │ results + logs v [Reviewer tab] │ v [Human approval] ─────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────────

In that sense, Hedr fits naturally into an AI-assisted workflow where:

  • an agent may draft or implement changes,

  • a developer may inspect and refine them,

  • tests and local services need to remain visible, and

  • approval gates still belong to humans.

This is an important distinction. Good orchestration is not about surrendering control. It is about making control practical.

Where Hedr fits in a modern workflow

For individuals, Hedr can make solo development feel less scattered. For teams, it can support clearer boundaries around ownership, review, and execution. In either case, its strongest use is in environments where multiple threads of work are active at once and the cost of losing track is high.

Typical use cases might include:

  • working across several repositories without mixing contexts,

  • running AI-assisted implementation alongside review tasks,

  • keeping local runtime processes available while switching focus,

  • maintaining a cleaner handoff between experimentation and production-ready work.

That makes Hedr especially relevant to teams adopting AI agents in a serious way. Once agents become part of normal engineering practice, the problem shifts from “Can they help?” to “How do we manage this safely and sanely?” Hedr belongs squarely in the answer to the second question.

Hedr is not the sort of tool that eliminates complexity altogether. Software development remains software development, which means it will continue to include trade-offs, edge cases, mystery failures, and at least one issue that disappears the moment someone else comes over to look at it. What Hedr can do is reduce unnecessary disorder so the remaining complexity is the kind that actually deserves your attention.

Final thoughts

If AI coding tools are expanding what a developer can do, tools like Hedr are expanding how well that increased activity can be organized. That may sound less dramatic than automated code generation, but it solves a more durable problem. Power without structure is just a faster route to confusion.

So if you are beginning to explore Hedr, the best place to start is with that simple idea: it brings shape to modern development work. Not with great fanfare, not with impossible promises, but with the deeply underrated gift of making the chaos easier to steer.

And in a universe full of terminals, agents, logs, and questionable decisions made without drinking tea, that is a remarkably useful thing.

Friday, 21 August 2026

How you could use ChatGPT and Codex to automate the useful, repetitive and occasionally alarming parts of the working day

 For most of its history, your computer has been an extraordinarily expensive way of waiting for you to type something. It has sat there, humming faintly, doing absolutely nothing of consequence, in a manner not dissimilar to a very well-paid security guard.

That is starting to change. ChatGPT's desktop app can now do things in the background — a scheduler, quietly working away while your computer is switched on and you are, in theory, elsewhere living your life. And the more of your accounts and files you're willing to connect it to, the more it can actually do for you. Which raises the obvious question: how far are you willing to let that go, exactly, before it's read all your emails and given what it now knows that it's sorry, but it's afraid it can't let you turn the computer off tonight, Dave.

From conversation to routine

OpenAI's scheduled tasks can run recurring work in the background. Tasks can use connected tools, plugins and reusable skills, while the desktop app can also work with local projects. Their status and recent runs remain available for review, which is important because an invisible automation with no audit trail is essentially a small bureaucratic ghost. OpenAI's scheduled tasks documentation

I started using this as a lightweight operational layer across my daily work.

Instead of writing a large custom application for every internal process, we describe the desired outcome, give the task controlled access to the relevant sources and schedule it. The task can then gather information, apply rules, create an output and report what happened.

The significant part is not merely that it runs at 8am. Ordinary schedulers have been doing that for decades with the tireless charisma of a boiler timer.

The difference is that the scheduled task can interpret what it finds.

It can distinguish an urgent request from conversational noise, understand that two differently worded alerts describe the same underlying dependency problem, rank work by importance and produce something intended for a human rather than another machine.

At least, that is the goal. As we shall discover, intelligence still benefits from being told exactly what "the same problem" means.

The daily briefing machine

One of our scheduled tasks produces a daily work brief.

It gathers information from Slack, Jira, Confluence and Google Calendar, then combines that with current weather, tide and surf information. The result is not a raw dump of everything that has happened since yesterday. That would be less of a briefing and more of an administrative avalanche.

Instead, the task:

  • identifies important Slack messages and unanswered requests;

  • retrieves the full queue of actionable Jira tickets;

  • prioritises work by status and importance;

  • includes relevant meetings and possible clashes;

  • checks current surf conditions around Newquay;

  • writes the finished briefing to the Desktop; and

  • sends a short Slack notification when it is ready.

The surf forecast is not strictly required for software delivery, but neither is tea, and history has shown the danger of removing either from a functioning organisation.

This workflow closely resembles the "work chief of staff" and daily briefing patterns in OpenAI's current automation examples: combining messages, meetings and work systems into a focused plan rather than simply presenting more information. ChatGPT and Codex automation use cases



Looking after the end of the day

We also use a scheduled task for end-of-day housekeeping.

It removes only specifically named temporary files, checks how much time has been logged for the day and sends a private reminder if the total is below the expected threshold.

The emphasis here is on "only specifically named."

Giving an automated agent permission to tidy a computer without defining exact boundaries is the Sorcerer's Apprentice problem: enchant a broom to fetch water, and it will keep fetching water long after the room, the castle and possibly the surrounding county are underwater. Something will certainly happen, but later investigations may struggle to classify it as improvement."

Our cleanup task therefore has explicit directory, filename and safety rules. Before deleting anything, it must verify that every target matches those rules. It is allowed to be useful, but not imaginative.

These boundaries should be set.

  • a narrow filesystem boundary

  • safe handling of symlinks and canonical paths

  • reversible actions

  • clear stop conditions

  • proven test behaviour on harmless fixtures before real use

That distinction matters. OpenAI's scheduled-task guidance recommends testing prompts before scheduling them, reviewing the first few runs and adjusting the instructions, tools or cadence where necessary. Local tasks also require the machine and desktop app to remain available when they need local files. Scheduled tasks: management and local projects

The weekly writing assistant

A weekly task searches our recent Confluence and Jira activity for a useful public topic. Which is how this blog was written.

It selects something with broader value, removes internal and company specific information, then turns the underlying lesson into a polished article draft. The finished document is written to the Desktop and announced in Slack.

The important boundary is that it does not publish the article.

Automation is excellent at collecting material, establishing structure and producing a strong first draft. Publication still deserves a human decision, particularly when the source material began life inside company systems.

The task accelerates the journey from "we learned something useful this week" to "here is an article someone can review." It does not quietly declare itself Head of Communications and begin issuing opinions on behalf of the organisation.

Turning security alerts into Jira work

Another automation monitors several Slack channels for vulnerability notifications from Snyk.

When a genuine new alert appears, the task extracts the package, installed version, severity, repository and available vulnerability information. It searches Jira for an existing match, creates a properly structured security task when necessary, and replies to the Slack alert with the resulting Jira link.

This is considerably better than relying on someone to notice a bot message while discussing a deployment, eating lunch or attempting to discover why a CSS rule has declared war on Safari.

It has also taught me one of the most valuable lessons in automation: the first version of a rule is always too literal. Rule do need tweaks over a few attempts.

That is where these tools become genuinely useful. An automation does not have to remain a brittle script forever. Its reasoning rules can be reviewed and refined when reality provides an edge case, which reality generally does with considerable enthusiasm.

Skills, plugins and connected systems

Scheduling provides the clock, but integrations provide the hands.

Plugins and connected tools allow ChatGPT and Codex workflows to work with services such as Slack, Jira, Confluence and calendars. Skills provide reusable operating instructions for particular kinds of work, helping the task apply the same process consistently across multiple runs.

OpenAI's documentation describes scheduled tasks as being combinable with skills for more complex work, and its current workflow catalogue includes bug triage, Slack prioritisation, verified operations, meeting follow-ups and continuously updated dashboards. OpenAI automation workflows

This means we can define not merely when something happens, but how it should be done:

  • where information should come from;

  • how duplicates should be detected;

  • which actions are permitted;

  • what requires human approval;

  • how success is verified;

  • what should happen when a source is unavailable; and

  • what evidence must be retained.

That last point is particularly important. A task should not say "everything was fine" because a search returned nothing. It should know whether the source was actually available, whether all pages were read and whether the result was genuinely empty.

There is a meaningful difference between "nothing happened" and "I failed to look." Humans have been exploiting this distinction in status meetings for generations.

What we have learned

The most successful automations share a few characteristics.

  1. They have narrow responsibilities. "Monitor these five Slack channels for Snyk alerts" is better than "look after security."

  2. They preserve evidence. Slack timestamps, source links, Jira keys and run records make it possible to understand what happened later.

  3. They are idempotent. Running the same task twice should not create the same ticket twice, send the same message twice or remove anything twice that ought to exist once.

  4. They fail visibly. If one connected source cannot be read, the task should not advance its checkpoint and quietly forget the missing interval.

  5. They keep humans in the appropriate part of the loop. Machines are good at repetition, comparison, collation and the relentless application of carefully written rules. Humans remain useful for judgment, accountability, changing priorities and recognising when three technically different vulnerabilities are, in practical terms, one dependency upgrade wearing several hats.

A quieter kind of automation

The aim is not to build an enormous autonomous system that runs the company while everyone retreats to the beach.




Although, to be completely honest, the surf report suggests the idea has received some preliminary consideration.

The real opportunity is quieter. It is to remove dozens of small acts of remembering:

  • check the security channels;

  • review the ticket queue;

  • assemble the morning brief;

  • clean up temporary files;

  • check today's time logs;

  • find a useful subject for the weekly article; and

  • notify the right person when something is ready.

None of these tasks is individually revolutionary. Together, however, they consume attention — the one organisational resource for which nobody has yet found a reliable package upgrade.

ChatGPT and Codex scheduled tasks give us a practical way to return some of that attention. They can watch, gather, compare, draft, create and report while leaving decisions and accountability where they belong.

The machine now has a diary.

Our job is to make sure we write very good instructions in it.

Friday, 14 August 2026

When an Email Says “Delivered,” What Has Actually Happened?

 A PRACTICAL GUIDE TO TRUSTWORTHY EMAIL STATUS

A useful guide to status messages, privacy boundaries, and resisting the universal human urge to call a tracking pixel a witness.

 




Email status can sound wonderfully conclusive. “Delivered” has the confident ring of a parcel landing on a doormat, saluting, and making a small speech. In reality, it means something more modest—and much more useful when explained clearly. 

If you send email for a product, community, or organisation, a simple status view can build trust. The trick is to make it accurate, privacy-conscious, and no more nosy than the job requires.

Use status words that mean what they say

A good status system separates the stages of a message's journey, so it can explain what actually happened without accidentally promising telepathy.

Accepted means the sending service took responsibility for the message — it left the building, so to speak. Delivered means the recipient's mail system accepted it, which is as far as the postal analogy can responsibly go: nobody signed for it, nobody read it aloud at breakfast.
Temporary failure means the journey has paused — the message equivalent of standing by the road with a thumb out, waiting for another mail server that happens to be going the same way. It may yet get picked up. Permanent failure means it's given up thumbing altogether and gone home: it isn't going anywhere until something changes, and no amount of patience out on the hard shoulder will fix that on its own.

"Opened" and "clicked" live in an entirely different category: engagement signals, not proof of anything. They're useful clues, but they don't confirm that a person read the message, understood it, or nodded along in agreement. Images get blocked, security scanners dutifully click every link in sight, and small robots have been known to be very enthusiastic readers indeed — enthusiastic enough that a suspiciously high open rate usually says more about the scanner than the recipient.


Build an event trail, not a single magic label

One email can have several meaningful events: it was accepted, retried, delivered, and perhaps later opened. Keep those events as a history rather than replacing yesterday’s fact with today’s headline.

This has two benefits. First, it gives support teams a truthful timeline when someone asks what happened. Second, it lets a simple display choose a sensible current state without destroying the evidence underneath it.

Give each outbound message a harmless unique reference, then connect incoming delivery events to that reference. Make the receiver tolerant of repeats: networked systems often deliver the same update more than once, because the universe prefers backups to confidence.

Make privacy the shape of the system

People should be able to see the status of their own messages without gaining a telescope into everyone else's. That means permission comes first, and the data view follows — a rule so obvious it barely needs saying, and so easily ignored that it needs saying anyway.

Access should be derived from the person actually signed in, not from whatever identifier they happened to type into a box — a distinction that sounds pedantic until you remember that boxes will believe absolutely anything you tell them. That same authorised boundary should apply everywhere a query might wander: reports, dashboards, the support tool someone built in an afternoon and never mentioned again. Collect only what helps, keep recipient data around only as long as troubleshooting genuinely requires it, and then let it go — data has a way of becoming a liability in exact proportion to how interesting it looked at the time. And before any delivery update earns a place in the record, make sure it really did come from the sending service, and not from something merely wearing its coat.

For many teams, the best first version of all this is not a grand control room with blinking lights, tempting as blinking lights are. It's a short periodic summary and a restricted lookup for the occasional question.


If acknowledgement matters, ask for acknowledgement

Sometimes delivery is not the outcome you need. If a process requires a person to confirm receipt, add a deliberate action: a secure confirmation link, a signed step, or an authenticated acknowledgement in the relevant service. Track that action separately from email delivery.

That small distinction makes communication both clearer and kinder. Recipients are not reduced to pixels; senders are not left reading tea leaves in an activity log.

The takeaway: report email delivery precisely, preserve the event history, and keep each person’s view tightly bounded. Clear language and careful access controls turn a mysterious status light into something genuinely trustworthy.