Journal

Three failures that said nothing

In 2026, the three failures that took most of my time on Jewely's sites had one thing in common. None of them showed an error. Everything appeared to work, and the result was wrong...

Article4 tagged threads

In 2026, the three failures that took most of my time on Jewely's sites had one thing in common. None of them showed an error. Everything appeared to work, and the result was wrong.

I tell them in order, then the rule I took from them.

1. The prices that stopped moving

February 2026, on Auberi's PrestaShop store. Prices come from the Odeis ERP: it drops a file, and the store imports it seven times a day. For several weeks, prices had not changed on the site.

I checked from the nearest cause to the farthest. The code in production was the versioned code. The import module computed and saved prices correctly. The scheduled task ran at the expected times. That left the file itself, and the transfer log answered:

articles.txt    133 bytes received
dispo.txt       245723 bytes received

The articles file weighed 133 bytes: a header line, an end marker, no products. The import read it, found nothing to update, and stopped without reporting anything. The stock file kept arriving normally, which hid the problem.

The cause was in the ERP export, on the client's side. It was not mine to fix. On my side, I entered the prices by hand from the client's file until the export was restored. Then I added the missing check, in the import dashboard:

$smallFiles = array_filter($articleFiles, function ($f) { return filesize($f) < 200; });
if (count($smallFiles) === count($articleFiles)) {
    $stats['health'][] = ['level' => 'danger', 'message' => 'Tous les fichiers articles sont vides (< 200 octets) — export ODEIS probablement en erreur'];
}

The message reads: "All article files are empty (under 200 bytes) — the ODEIS export is probably failing."

An import that receives zero lines is not a successful import. It is a question to put to someone.

2. The deployment reported as successful

2 March 2026, on Jewely's Docker Swarm cluster. The disks of two servers were more than 95% full. They could not download the new image, which weighed about 786 MB. Swarm went back to the old version, and the deployment command still returned a success code.

The incident is told in detail in When a deployment succeeds without really succeeding.

What I wrote afterwards is a deployment agent. It does not believe the command. It compares the digest of the expected image with the one actually running, and it looks at the disks before starting:

What the agent finds What it does
Disk above 80% Targeted cleanup of old images, then deployment
Disk above 90% Deployment blocked
The running image is not the expected image New attempt, three at most

This agent checks beside the pipeline. Putting the same check inside the automatic trigger was up to the administrator of the AWS servers, and the documentation says so.

3. The cache that could no longer be cleared

March 2026, in the common code. The site's cache is organised by tags, which makes it possible to clear everything about a product or a page in one go. Tags only exist with Redis. In development and in tests, the cache is a plain file, without tags.

A first version worked around the gap: without tags, it still wrote to the cache, under another key. Writing worked. Clearing by tag did nothing. A cached value could therefore never be invalidated again, and it stayed wrong without any message.

The decision, written in an architecture note: without tags, no cache at all.

public static function rememberWithTags(array $tags, string $key, int $ttl, callable $callback): mixed
{
    if (self::cacheSupportsTagging()) {
        return Cache::tags($tags)->remember($key, $ttl, $callback);
    }
    return $callback(); // bypass cache entirely
}

The cost: in development, every call goes to the database. I accepted it. In that place, a slow and correct answer is worth more than a fast and stale one.

The same defect, three times

What the system showed What was happening The check added
Prices Import finished No article read Alert when every file is empty
Deployment Success Old version in production Comparison between the expected image and the running one
Cache Value served Value never invalidated No cache when it cannot be cleared

In all three cases, someone had planned a quiet way out of an awkward case. Zero lines: nothing to do. Image that cannot start: go back to the old one. No tags: write somewhere else. Each way out was reasonable on its own. None of them told anyone.

The rule

A failure you can see is worth more than a fallback that stays quiet.

I then applied it to the CMS templates: when a house is missing a template, the site raises an error. It no longer goes looking for another house's. A missing template shows up in pre-production. Another house's template would have gone to production.

Before putting an automated flow into service, I now ask four questions:

  1. What do you see when nothing arrives? An empty file, an empty list, zero lines processed.
  2. Does "success" describe the command, or the state reached? A command that returns 0 does not say which version is running.
  3. Is there a fallback? If so, what does it serve instead, and who finds out?
  4. What does the noise cost? One alert too many costs a minute. Wrong prices for weeks cost something else.

The pipeline and its hardening are described in the case study Checking what really runs after an ecommerce deployment.