Skip to content
Conversion & marketing

Content audit: cleaning up old website content

Old pages, duplicate topics, outdated prices: how a content audit sorts an existing inventory — with four exits, fixed criteria at the entrance and redirects.

15 min read Content-AuditContent-MarketingWebsite-Pflege

Any website that has been running for a few years carries pages nobody remembers: the campaign page from 2021, three variants of the same service description, an article quoting a price that no longer exists. The inventory grows because adding is easier than deciding. A content audit turns that around. It turns the inventory into a list and sends every row through exactly four exits: keep, revise, merge, redirect. That such an inventory is the normal case rather than a failure is clear from the numbers. 79.01 percent (Eurostat) of EU enterprises with ten or more employees run a website, and 66.36 percent (Eurostat) use it to describe goods, services and prices — the part that reality overtakes fastest. This article shows how the inventory is built, which criteria apply at the entrance, how to merge and redirect without giving away visibility, and what rhythm makes the effort sustainable.

Key takeaways

  • An audit starts with a list, not with a decision: one row per address, plus views, enquiries, the last substantive edit and the reason the page exists.
  • There are four exits and no fifth one. Leaving a row blank means you have kept the page without deciding to keep it.
  • Deleting in bulk so the site looks fresher achieves nothing: the documentation asks exactly that question and answers it in brackets with “No, it won’t” (Google Search Central).
  • Merging means redirecting. Status codes 301 and 308 (Google Search Central) are the permanent signal, and 68 percent (Web Almanac 2025) of desktop pages already carry a canonical tag.
  • Capacity is the scarce resource: 39 percent (Content Marketing Institute) of the marketers surveyed list time, people and budget among their three biggest challenges. An audit therefore prioritises instead of touching everything.

A content audit is an inventory, not a deletion round

The term sounds like an exam with a grade attached; what it means is something more sober. You write down what exists, measure it against a small set of fixed criteria and decide, page by page, what happens next. The value sits in the decision, not in the analysis. Count only, and you end up with a spreadsheet; sort, and you end up with a shorter, more current website. The starting position is much the same in almost every inventory: things were added for years because adding needs no agreement, and rarely removed because removing needs a justification.

The inventory itself is the rule, not an oversight. In the EU, 79.01 percent (Eurostat) of enterprises with ten or more employees had a website, 76.66 percent (Eurostat) among small enterprises and 95.65 percent (Eurostat) among large ones. Most of what sits on those sites is service description and price information: 66.36 percent (Eurostat) of enterprises publish exactly that, with job adverts following at 31.34 percent (Eurostat). That also tells you which part of the inventory ages first — a price, an opening time, a name in the team section.

What stands out is less the absence of plans than the absence of feedback loops. 97 percent (Content Marketing Institute) of the B2B marketers surveyed say they have a content strategy, yet only 59 percent (Content Marketing Institute) consider their own work at least somewhat effective; the survey covers 1,015 respondents, mostly in North America. Between the plan and the effect sits the inventory: content that was once produced to plan and has not been looked at since. What a plan that builds in maintenance from the start looks like is covered in our article on content marketing strategy for smaller companies.

Tidying up is not thinning out

It is tempting to turn an audit into a deletion round: fewer pages, a clearer picture. The quality documentation explicitly asks whether a lot of older content is being removed “primarily because you believe it will help your search rankings overall” — and answers itself in the same sentence with “(No, it won’t)” (Google Search Central). So what gets removed is what should go for editorial reasons, not what might flatter the statistics.

What sits in a grown legacy inventory

Before anything is sorted, you need the list, and it is usually easier to obtain than the inventory suggests. More than 54 percent (Web Almanac 2025) of observed websites run on a content management system, and roughly 64 percent (Web Almanac 2025) of those run on the single most widespread one. In practice that means there is a database, there is an export, and the rows for your inventory do not have to be typed out by hand. Where no export exists, a crawl of your own domain produces the same columns, only more slowly.

How much text accumulates per page is sobering and helpful at the same time. The median inner page carries 339 words (Web Almanac 2025) on desktop and 323 words (Web Almanac 2025) on mobile. The average blog article sits well above that at 1,333 words (Orbit Media), but it is the exception in the inventory. Most of the work therefore concerns short pages, where the decision comes quickly: either the page carries its statement on its own, or it belongs to another one.

The technical columns fill themselves almost automatically, because the typical gaps look the same everywhere. A title element is present on practically every page, namely on 98.6 percent (Web Almanac 2025) of desktop pages. The meta description, by contrast, is missing from almost one page in three; it is set on 67.7 percent (Web Almanac 2025) of desktop pages and 67.2 percent (Web Almanac 2025) of mobile pages. And where a title exists, it is often too long: the median sits at 77 characters (Web Almanac 2025) on desktop and 79 characters (Web Almanac 2025) on mobile. Three columns, in other words, that can be filled by machine and worked through in batches — further candidates are listed in the technical SEO checklist.

Address and title

The full address, the current title and the meta description. Three columns that any content management system can export.

Views and entries

Page views over the last twelve months and how many of them were entry points. Without that figure, sorting runs on gut feeling.

Enquiries and deals

Which enquiry did this page trigger? A page with few views but two enquiries a year is no candidate for deletion.

Last substantive edit

Not the year in the footer, but the day someone last changed a statement. That difference decides the exit later on.

Overlap

Which other page answers the same question? This column decides the merge exit and is the one most often skipped when filling in the sheet.

Legal status

Prices, deadlines, mandatory details and statements that follow a rule. They age regardless of how many people open the page.

  • The page export from the content management system delivers address, title, date and author in one go.
  • Analytics deliver views, entries and time on page per address — without them the inventory is an opinion poll.
  • The sitemap shows which addresses are submitted for indexing at all, and exposes the first gap.
  • A crawl of your own domain finds pages that hang in no navigation any more and are still reachable.
  • The enquiry form or the customer record shows which page a request came from.
  • The site's internal search shows what visitors look for and where they come up empty.

The four exits of the sorting line

Every row gets exactly one exit. That sounds like bureaucracy and is the actual trick. As soon as a fifth option is allowed — “look at again later” — half the inventory moves there, and the audit ends as a spreadsheet nobody opens again. Four exits, one decision per row, one responsible person per exit: the sorting needs no more rules than that.

Keep means the statement holds, the figures are current and the page is in demand; at most it gets fresh links. Revise means the topic works but the text does not. Merge means two or three pages say the same thing: one stays and inherits the best parts of the others. Redirect means the offer is gone while the address very much is not — it has inbound links, bookmarks and a history in search.

ExitWhen it appliesWhat to doHow success shows up
KeepStatement holds, demand is thereRefresh links and dates, nothing elseViews stay stable, no rework needed
ReviseTopic works, text is thin or datedShorten the title, set the description, renew the evidenceEntries and time on page rise at the same address
MergeTwo pages answer the same questionMove the best paragraphs over, set canonical, redirectOne address collects the views of both predecessors
RedirectOffer is gone, address has inbound links301 to the closest remaining page, rewire internal linksNo error pages in the report, views land on the target

For smaller inventories this is a question of clarity, not of technology. The crawl budget documentation calls itself an advanced guide aimed primarily at large sites: from around 1 million pages (Google Search Central) with weekly changes, from 10,000 pages (Google Search Central) with daily changes, and inventories where a large share of the URLs sit in Search Console under “Discovered - currently not indexed”. The same source calls those numbers a rough estimate for classifying your own site and expressly not exact thresholds. One piece of advice from it carries over to small inventories as well: eliminating duplicate content focuses crawling on unique content rather than unique URLs (Google Search Central). Run three pages on the same topic and you split links and attention three ways instead of one — an effect that is especially visible with location pages for multiple sites.

The most expensive exit is the one nobody takes

Deleting without a redirect feels fastest and hurts longest. Inbound links run into nothing, bookmarks do the same, and the page disappears from the index without passing its effect to another address. So the rule is: when a page disappears, it gets a successor — the closest remaining page that answers the same question. Only when there is none does the address end with a clean 404 or 410.

The criteria at the entrance

At the entrance to the sorting line stand questions that can be answered with yes or no. Anything that starts a discussion does not belong in a criterion but in the notes column. One order has proven itself: effect first, then truth, then technology. A page without views and without enquiries is a candidate but not a verdict — it may be the only page answering a particular question and may simply have gone unlinked.

A date is only a criterion if it means something. The value in the sitemap should carry the last significant update to the page (Google Search Central) rather than the day of the last deployment. By the same token, the year in the footer proves nothing about whether anyone has read the statement above it. Where prices, discounts or deadlines appear, the check gets stricter: discount claims follow rules of their own, as the article on the 30-day rule for strike-through prices explains in detail.

inventory-row.json
{
  "address": "https://www.company.example/services/maintenance-contract/",
  "title": "Maintenance contract - Company Ltd - request now",
  "views_12m": 412,
  "entries_12m": 96,
  "enquiries_12m": 3,
  "last_content_change": "2022-03-14",
  "duplicate_of": "/services/maintenance-and-service/",
  "legal_status": "price from 2022, deadline unclear",
  "exit": "merge",
  "reason": "same question, fewer views, outdated price"
}

A row like this carries the whole decision. It says how the page is found, what it produced, when someone last wrote into it, what it duplicates and where it goes. Leave the reason column blank and you will miss it three months later: every follow-up question puts the decision back on the table. The same structure works as a handover to everyone who maintains the inventory later on — more on that on our page about content marketing.

  1. Does the page bring views, entries or enquiries? If so, it is checked rather than touched.
  2. Are prices, deadlines, names and mandatory details still correct? Wrong weighs heavier than thin.
  3. Does another page answer the same question better? Then the merge exit is settled.
  4. Does the offer still exist? If not, the address is a case for the redirect exit.
  5. Does the page carry a title, description, canonical tag and alt texts? Those four points cost minutes.
  6. Is there an inbound link to the page? Then it gets redirected rather than removed.

Turning the date over is not an update

The quality documentation explicitly asks whether page dates are being changed to make pages seem fresh when the content has not substantially changed (Google Search Central). An audit that leaves behind two hundred new dates and not a single changed statement has not improved the inventory; it has only covered its tracks. The date is the result of the work, not a substitute for it.

Merging and redirecting without giving away visibility

Two tools are involved in merging, and they do different things. The canonical tag says which of several near-identical addresses is the authoritative one; it sits on 68 percent (Web Almanac 2025) of desktop pages and 67 percent (Web Almanac 2025) of mobile pages, making it the normal case. Actually being pointed at another address is rare by comparison: only 7 percent (Web Almanac 2025) of desktop pages and 9 percent (Web Almanac 2025) of mobile pages are canonicalised. Merging is therefore the uncommon move in an inventory — and precisely for that reason the one that shifts the most.

The redirect is the stronger signal. Status codes 301 and 308 (Google Search Central) mean that a page has permanently moved to a new location, and they are treated accordingly. Reaching for the removals tool instead buys a deadline: requests there last about 6 months (Google Search Central), after which what the page itself serves counts again. A noindex directive is the exception too, appearing on 3.5 percent (Web Almanac 2025) of desktop sites and 2.4 percent (Web Almanac 2025) of mobile sites. How a larger rebuild runs without losing rankings is covered in the article on the most common website relaunch mistakes.

Terminal
$ curl -sI https://www.company.example/services/maintenance-and-service/ | head -2
HTTP/2 301 location: https://www.company.example/services/maintenance-contract/
$ curl -s https://www.company.example/services/maintenance-contract/ | grep -o 'rel="canonical"[^>]*'
rel="canonical" href="https://www.company.example/services/maintenance-contract/"
$ curl -s https://www.company.example/sitemap.xml | grep -c '<loc>'
166

After the merge comes the follow-up work, and that is the part that tends to be left lying around. A sitemap holds at most 50,000 addresses (Google Search Central) per file, which is ample for most inventories — but it has to gain the new targets and lose the old ones. Internal links have to come along: the median page carries six outbound links (Web Almanac 2025), so the stock of links is small enough to walk through completely. And if something still fails to take effect afterwards, look inside the head: 10.1 percent (Web Almanac 2025) of desktop sites carry an invalid element there, behind which everything else loses its effect.

Two special cases deserve their own attention. Language versions hang on hreflang, which 20.3 percent (Web Almanac 2025) of desktop pages carry; merge one version without pulling the other's references along and you leave a chain pointing into nothing — the article on multilingual websites and hreflang sets out the order of work. And structured data is served by 50 percent (Web Almanac 2025) of home pages; markup pointing at a merged address shows up in testing tools straight away.

What else surfaces during the clean-up

An audit that opens every page anyway should pick up three more columns, because they come at no extra cost. The first is images: 13 to 14 percent (Web Almanac 2025) have no alt attribute, another 30 percent (Web Almanac 2025) use an empty one, and 8.5 percent (Web Almanac 2025) of alt texts simply repeat the file name with its extension. That last pattern can be found by machine and fixed in batches; what images should look like afterwards is covered in the article on optimising website images.

The second column is colours and contrast that have grown over the years. Only 31 percent (Web Almanac 2025) of mobile sites meet minimum colour contrast requirements, and the median accessibility score sits at 85 percent (Web Almanac 2025) — a figure that has barely moved for years. Old pages with washed-out greys are the most common find, and the correction is usually one line in the stylesheet. The article on accessibility in practice describes what to watch for during that rework.

The third column is weight. The median home page grew 7.8 percent year over year to 2.7 MB (Web Almanac 2025); on mobile it is 2,362 KB (Web Almanac 2025) against 845 KB (Web Almanac 2025) ten years earlier. Home pages carry 239 percent (Web Almanac 2025) of the image bytes of comparable inner pages. Old hero images that were long since replaced and are still being loaded are a typical find. And before pages are taken out of service, robots.txt belongs on the list: it is missing entirely on 13.3 percent (Web Almanac 2025) of desktop sites.

Some content ages regardless of how many people read it. The general information duties for commercial digital services are set out in section 5 of the German Digital Services Act; the details have to be easily recognisable, directly accessible and permanently available (Digitale-Dienste-Gesetz), and the provision lists eight numbered items (Digitale-Dienste-Gesetz) for that purpose. In an audit this is a single yes-or-no row, and it still turns up regularly: legal form changed, register entry moved, supervisory authority renamed. The article on legal notice requirements works through the points one by one.

Accessibility comes on top of that. The German Accessibility Act applies to services provided to consumers after 28 June 2025 (Barrierefreiheitsstärkungsgesetz). Micro-enterprises are exempt, defined as companies with fewer than ten people and at most 2 million euros (Barrierefreiheitsstärkungsgesetz) in annual turnover or balance sheet total. For services delivered with older products, a transitional provision runs until 27 June 2030 (Barrierefreiheitsstärkungsgesetz). Breaches carry fines, with the range reaching up to 100,000 euros (Barrierefreiheitsstärkungsgesetz) depending on the case. What that means for a grown website is covered in the article on the obligation to run an accessible website.

Notwithstanding subsection 2, service providers may continue until 27 June 2030 to provide their services using products which they lawfully used before 28 June 2025 to provide the same or similar services.

Section 38(1) of the German Accessibility Act, translation of the official wording

Rhythm, effort and priority

An audit rarely fails on method and often fails on the calendar. 39 percent (Content Marketing Institute) of the marketers surveyed list time, people and budget among their three biggest challenges, 40 percent (Content Marketing Institute) name content that prompts a desired action, and 33 percent (Content Marketing Institute) name measuring effectiveness. That is exactly why the order matters more than completeness: pages with enquiries first, then pages with views but no enquiries, then the silent remainder. Work the other way round and the momentum is spent on pages nobody opens.

How sober the picture is shows in the self-assessment: only 21 percent (Orbit Media) of 808 respondents report strong results. An inventory that runs through the sorting line completely once a year, with quarterly touch-ups on the pages that bring enquiries, is more realistic than a one-off push. Which metrics are worth watching is covered in the article on measuring website success; anyone who cannot hold that rhythm in-house will find it again in ongoing website care.

The inventory is a line, not a project

An audit planned as a one-off project ends with a tidy list and an inventory that starts growing again the following month. It only takes effect as a recurring pass with fixed dates, fixed criteria and one responsible person per exit. The structure that comes out of it carries further: into editorial planning, into search engine optimisation and into every conversation about whether a new page really needs to be a new page.

Sources and studies

This article is based on data from: The 2025 Web Almanac (HTTP Archive, chapters on SEO, Accessibility, CMS and Page Weight), Google Search Central (sitemaps, crawl budget, redirects, removals tool and quality documentation), Eurostat (Digital economy and society statistics – enterprises), Content Marketing Institute (B2B Content and Marketing Trends), Orbit Media (Annual Blogger Survey), the German Accessibility Act and the German Digital Services Act.

Related Articles

Conversion & marketing

Content Marketing Strategy for SMBs: the Plan

How small and medium businesses build organic visibility over time with topic clusters, an editorial calendar, SEO texts and honest measurement.

12 min read
Law, privacy & accessibility

Who Really Owns Your Website: Domain, Access, Rights

Domain, credentials and rights of use are three separate levels. Who is entered as the domain holder, how a change of provider works and what belongs in the.

16 min read
Online shop & e-commerce

Christmas Trading 2026: Prepare Your Online Shop Now

A dated plan for Christmas trading: six stages from calendar week 38 to the returns window — product data, mandatory information, load testing, emergency plan.

15 min read