At 11:00 one morning, my fitness coach generated a fresh workout for a man who had visibly not done the 08:00 one.

It was a good workout. Three tiers, sensible progression, correctly cautious about a bad shoulder. It was also the second one that day, which meant I now had two workouts I was not doing instead of one.

That is the whole problem with putting a language model on a schedule, and it took me about three weeks to see it properly.

The easiest automation, and the one nobody tunes

Scheduling a prompt is close to free. You write the instructions once, choose a cadence, and the model runs whether you are awake or not. No server, no queue, no cron file, no integration platform sitting in the middle charging you per task. Most of the assistant products ship some version of this now, and in my experience most people never switch it on.

I switched it on five times. A nightly intelligence brief, a fitness coach, a weekly review, a market monitor, and a daily job hunt. Between them, roughly six thousand words of prompt, written mostly at night over three weeks.

All five worked immediately. All five were quietly useless within a month, and the reason was the same every time.

They cannot tell you that nothing happened

Ask for a market briefing every morning and you will get a market briefing every morning. Not because the market did anything. Because you asked.

This sounds obvious written down, and it is the root of nearly every complaint about these things being noisy. A model handed a reporting job completes the reporting job. Nothing in the setup gives it a reason to come back and say “nothing occurred, go back to sleep,” so it doesn’t, so it invents a reason to be interesting instead.

What you get is an automation that technically works and that you stop reading inside three weeks. That is a hell of a result for something you built to save yourself attention. A briefing you skim costs the same as one you read, and it trains you to ignore the channel it arrives on, which is the channel you will later want for something that matters.

So the design problem was never prompt quality. Getting good output is trivial. Getting deliberate silence is the entire job.

The sentences that fixed it

I went back and read my own prompts, and the useful discovery was that the load-bearing parts were not the parts I was proud of. They were the clauses I bolted on later, in the edit modal, annoyed, after something had gone wrong.

If there is no meaningful new development, do not notify me.

Do not duplicate an alert from the market monitor unless something materially changed after that alert.

Every one of those is scar tissue, and each one names a specific failure. The monitor was waking me up about an event it had already reported, because a second outlet had written a second article about the same announcement. The nightly brief was faithfully repeating the monitor back to me an hour later. Neither is a model being stupid; both are me failing to define an interface between two tasks that overlapped.

Three fixes came out of that, and they generalise past my setup.

Define the unit as the event, not the article. Five outlets covering one announcement is one thing that happened. Re-alert only on a new fact, an official confirmation, a changed number, an escalation. Otherwise your alert volume tracks press coverage rather than reality, which is a strange thing to have built on purpose.

Name the owner when scopes overlap. Once you have more than two of these, they will collide. Mine now says explicitly that the job hunt owns job discovery and the nightly brief must not surface individual postings. That is not prompt engineering, it is the same boundary-drawing you would do between two services.

Split the producer from the checker. This is the one that fixed the fitness coach. The 08:00 run produces the session, and the 11:00 run does exactly one thing: it checks whether I said anything to it today, and if I did, it exits without a sound.

A reminder fires on a clock. A checker fires on a clock and then tests a condition, and passing that test means I never hear from it. The second run must never regenerate the first run’s output, or you have not built a checker, you have built a second producer and a guilt machine.

The one that is actually hard

There is a constraint underneath those three that quietly invalidates them, and I had it wrong in all five tasks.

Most scheduled tasks begin each run with no memory of the previous one. Which means “report only what changed since last time” is an instruction the model cannot verify. It has nothing to compare against. It will comply anyway, confidently, because that is what it does.

My nightly brief teaches me one new word of a language I am learning, and was supposed to build on the day before. It depended entirely on the model scrolling back far enough in its own conversation to find yesterday’s word. Sometimes it did. When it didn’t, I got Monday’s word again on Thursday and had no way to know, because the output looked identical either way.

The fix is to give the task somewhere durable to write its conclusions and have the next run read that first. Worth knowing before you do it: once a thread or a file holds the accumulated state, deleting it deletes the memory. What used to be a disposable log becomes load-bearing, and nobody tells you the day that changes.

Kill criteria, or you are just collecting

The last thing I added to all five was a review date, written into the prompt text itself, so each task carries its own expiry. If it has not changed a decision in thirty days, it gets deleted rather than paused.

I added that because I went looking at my other scheduled jobs on a different system and found three, of which exactly one was alive. The other two were one-shot tasks from two weeks earlier, finished, sitting in the list looking exactly like the working one. Nothing in that list distinguished a job doing work from a job that had already finished, which is its own version of the same bug: a display that cannot express the empty case.

That is the generalisation, if there is one. A counter that cannot report zero is not a counter. A status line that prints the same value busy or idle is not a status line. A briefing that always has news is that identical defect in prose, and it is more dangerous than the numeric kind because fluent text reads like judgment.

Where I would draw the line

One practical note, since this is where I see people waste the most time. Anything that touches your repositories, your builds, or the files on your machines belongs in tooling that actually has access to them. Scheduled prompts are good at watching things outside your walls: a market, a job board, a public status page, a competitor’s changelog. They are bad at being the thing that knows the state of your own systems, because they mostly cannot see them, and a task that cannot see is a task that guesses.

None of the seven or eight constraints above are exotic. They are what on-call alerting figured out decades ago, rediscovered in a place where setup takes a text box and nobody is thinking about alert fatigue yet.

The question I now ask of any scheduled task is not whether its output is any good. It is what would have to be true for this to send me nothing, and whether that has ever once happened.