A page that is not in the index does not rank badly — it does not rank at all. This article shows why crawl budget, discovery and indexing are three separate stages, and how the Indexing Hub in the new Semalt Panel makes that blind spot measurable.
Most SEO projects in Vienna and the surrounding region start with content, keywords and links. That is understandable: it is where the visible work happens. What gets overlooked is the stage before it — whether a search engine knows those pages at all and has accepted them. While that question is open, every further optimisation is work on an object that does not exist as far as search is concerned.
The problem is not confined to large portals. An online shop with 8,000 product pages, an estate agency with thousands of listings, a specialist publisher with a grown archive — all accumulate pages faster than those pages get picked up. Because indexing stays invisible day to day, the shortfall only surfaces when organic traffic has been flat for months despite steady editorial output. For the wider picture, the overview page for the new Semalt Panel covers the platform; this article stays with the indexing chapter.
Crawl budget, discovery and indexing are not the same thing
The most common mistake in indexing questions is assuming this is one process. In reality there are three consecutive stages, each with its own logic. A problem at the first cannot be repaired by measures taken at the third — and the reverse holds just as firmly.
Discovering that an address exists, retrieving it, and deciding whether to accept it into the index are separate steps. Keeping that separation in mind means that when visibility drops you ask the right first question: which stage is failing?
| Stage | What happens there | Typical blockage | The right countermeasure |
|---|---|---|---|
| Discovery | The search engine learns a URL exists — via internal links, sitemaps, external references or an active submission. | Orphaned page, missing sitemap entry, click depth beyond ten levels. | Internal linking, a clean sitemap, active submission. |
| Crawling | The bot actually fetches the address. Fetch volume per period is what people loosely call crawl budget. | Slow server response, 5xx under load, millions of worthless parameter variants. | Infrastructure, status codes, tidying URL variants. |
| Indexing | The retrieved page is assessed and accepted — or not. An editorial decision, not a mechanical one. | Thin content, duplicate, no standalone value, contradictory canonicalisation. | Substantive content, clear canonicals, consistent signals. |
Crawl budget is not a figure on an invoice. It is the outcome of server performance, past response quality and an assessment of how worthwhile a domain is. A server that errors during traffic peaks, or takes three seconds to the first byte, gets fewer fetches — the search engine throttles itself so as not to strain the site.
Six reasons why pages stay out
Across audits the causes repeat with remarkable reliability. The six patterns below cover most of what turns up in practice, sorted by the stage at which each takes effect.
Orphaned pages
A URL exists and opens fine, but no other page on the site links to it. It has no route in at all.
- Arises in category rebuilds and removed filters
- Campaign pages pulled from navigation after the promotion
- Stays online, loses every form of findability
Weak internal linking
Subtler than the orphan: formally linked, practically unreachable. Click depth reads as a relevance signal.
- A product on page 14 of a pagination
- Twelve clicks from the home page
- Effectively invisible to users and bots alike
Duplicate content
Shops generate duplicates almost by themselves. The search engine then picks a variant — not necessarily yours.
- The same product under three category paths
- Sorting, filter and session parameters
- Print views, variants with and without a trailing slash
Broken canonicals
The canonical tag is the right instrument against duplicates, yet it is often turned against itself.
- Every pagination page points at page one
- The target redirects, carries noindex or is gone
- A canonical is a hint, not an instruction
robots.txt and noindex
Two mechanisms constantly confused with one another — with expensive consequences at launch.
- Blocked URLs are never fetched, so their noindex goes unread
- Removal from the index needs crawlability plus noindex
- A staging noindex left in after a relaunch costs weeks
Slow delivery
Response times cap the number of fetches per unit of time. Irrelevant on small sites, decisive on large ones.
- 200 milliseconds against 2 seconds is a multiple in throughput
- At 30,000 URLs that decides days against months
- Server errors feed back into the whole domain
If you run a large page volume and have not got server response under control, start there, before any thought of submission tools. This sequence is the difference between a measure that works and a budget simply used up — the technical groundwork belongs at the beginning for good reason.
The Indexing Hub and its hard limits
The Indexing Hub is where submission, sitemap processing and observation come together in the panel. For planning, its limits are what matter: they determine what is realistic in a given period. Knowing those figures openly is worth more than any promise.
The most interesting figure is the smallest. A thousand URLs a day sounds generous for a company website with sixty pages. For a five-figure catalogue it is a budget to be planned like an advertising spend — and that is where the benefit lies: the limit forces a decision about which pages matter, one that was due anyway.
| Metric | Value | What it means in practice |
|---|---|---|
| Daily allowance, URL tracker | 1,000 URLs per day and account | The real throughput ceiling. A thousand chosen URLs beat a thousand arbitrary ones. |
| Bulk submission | up to 10,000 URLs per batch | A whole list in one step instead of manual slicing. Processing still follows the daily allowance. |
| Sitemap source | file upload or URL | Sitemaps not publicly linked yet can be checked ahead of a relaunch. |
| Recursive parsing | up to 3 levels deep | Index sitemaps pointing at further index sitemaps are resolved. Deeper nesting must be rebuilt flatter. |
| Sitemaps per job | up to 1,000 | Covers large sets segmented by category or language in one pass. |
| Concurrent jobs | maximum 2 | Two domains or two sitemap sets in parallel; the rest waits. |
| Queue | maximum 20 jobs | Enough for an agency with several clients, but not unlimited batch operation. |
| Per-URL log | timestamp, status, error detail | Shows whether an address was fetched, and with what result. |
Two ways in: batches and sitemaps
The hub offers two complementary entry points. One starts from a list of specific addresses, the other from an existing sitemap structure. Both feed the same tracker and share the same daily allowance.
Bulk submission of URLs
For when a very large number of addresses is new or changed at short notice.
- One step instead of many. The complete list is handed over once; splitting it into daily portions by hand is unnecessary.
- The daily allowance is still the brake. A batch of 10,000 addresses is processed over time within the 1,000 per day.
- Ordering is a decision. With a tight allowance, the sort order of your list determines the first week's commercial return.
- Worthwhile for exceptions, not as routine. A relaunch with a new URL structure, a migration, seasonal categories, a new range, a blockage just removed.
Sitemap processing
For established inventories already organised as an index sitemap structure.
- Upload or URL. The file can be uploaded or given as an address — useful for checking a new structure before it goes live.
- Recursive to three levels. An index sitemap pointing at further index sitemaps is resolved. Nest more deeply than that and you should rebuild it flatter.
- Up to 1,000 sitemaps per job. Sets segmented by language, location or content type run through in a single pass.
- Two jobs in parallel, twenty waiting. Enough for multi-client work, but set the order deliberately: the client relaunching on Friday goes before routine maintenance.
In both cases the submission goes out through the IndexNow interface to GoogleBot and BingBot. To see how the Indexing Hub built by Semalt fits together with the remaining modules, the platform overview describes it.
Push instead of waiting — and what the log says afterwards
IndexNow as an active notification
Classically, indexing works on a pull principle: the search engine decides when it comes round. On a rarely updated domain that can take weeks for pages deep in the structure. IndexNow reverses the direction — the site reports that something has changed at a given address.
The gain lies mainly in timing. A corrected product description, an updated price, a reworked guide page, a new category: instead of hoping the next scheduled visit picks the change up, the event is reported. For seasonal businesses that is the difference between visibility in time and visibility after the season.
Realistically, the push does two things. It shortens the time to discovery, and it replaces the unanswerable question of whether the search engine knows a page with the answerable one of why it decided against accepting it. Real progress, even if it sounds unspectacular.
The bot visit log as a diagnostic tool
Diagnostically, the most valuable part of the hub is not the submission but the per-URL log. For every address handed over, it records when a bot visited, what status that ended in and what error detail came with it. Live counters for submitted, found and failed addresses mean a running job is never watched blind.
The value comes from combining the three. The timestamp answers whether a fetch happened. The status separates technical cases from editorial ones. The error detail gives the direction for the fix.
- No visit despite submission. Reported, not fetched. Check robots.txt, server load and whether the domain is being throttled generally.
- Visit with a server error. The crawler arrived and hit a 5xx. An infrastructure matter affecting crawl behaviour across the domain, not just this page.
- Visit with 404 or 410. Usually a stale sitemap or an internal link to removed content. Tidy up rather than resubmit.
- Visit with a redirect. You submitted an intermediate step. Always hand over the redirect target instead.
- Successful visit, no acceptance. The technical part is done. The cause lies in the content, in canonicalisation or in missing internal links.
40,000 URLs at 1,000 per day
The scenario below is constructed to make the order of magnitude tangible. It is a worked example, not a result from a specific project.
A shop carries 40,000 addresses in its sitemap. At 1,000 a day that is 40 days on paper, provided the allowance is used in full daily and nothing needs resubmitting. Correction runs get added, so six to eight weeks is more realistic. The batch of up to 10,000 addresses does not shorten that — it only spares you the manual slicing.
Six weeks is too long for a seasonal push. The answer is not to reject the tool but to sort the list. A plausible breakdown of those same 40,000 addresses:
- Around 12,000 variants and parameter pages. Colour and size combinations, filters, sort orders. Canonicalise them, do not submit them. Allowance spent: none.
- Around 6,000 permanently unavailable items. They need a decision about redirecting or removing, not a submission.
- Around 400 category and guide pages. They carry the lion's share of revenue and go first — under half a day of allowance.
- Around 5,000 product pages with search volume and stock. They follow over five days and cover the running business.
- Around 16,000 long-tail pages. These continue in the background afterwards, with no deadline and no pressure of expectation.
Forty days of blind submission thus becomes barely six for everything that counts commercially. The difference lies not in the technology but in the ordering. Put the Search Console visibility data alongside it and you can align priorities with actual impressions too; the corresponding views on the same platform sit in the same account and need no separate export.
Sitemap hygiene and a workable routine
A sitemap is not an inventory of every address that exists; it is a recommendation. Each entry states: this page deserves indexing. If thirty per cent of the file is redirects, error pages and noindex addresses, trust in the whole file drops — and with it its value for the pages that genuinely belong there.
Indexable destination pages
Canonical addresses only, status 200, final spelling, correct protocol and host.
- Honest lastmod values, not dates reset daily
- Segmentation by category, product, editorial, location
- An index sitemap on top — it suits recursive processing
Everything non-indexable
Any address that will be rejected anyway dilutes the file's message and consumes attention.
- noindex pages, addresses blocked in robots.txt, redirects
- Parameter and filter variants, internal search results
- Basket, account area, staging hosts, error pages
From this follows a sequence that works in most projects, and it deliberately starts with tidying rather than submitting. First the sitemap is cleaned. Then it is processed — by upload or by address, depending on whether the file is public yet. The result gives a first picture: how many addresses were found, how many failed, where status codes cluster.
That error list is worked through before any allowance goes on submissions. Only then does prioritised handover begin, business-critical pages first. After a week the log holds enough data for the real analysis: visited and accepted, visited and discarded, never fetched. The third group points to technical causes, the second to editorial ones. Both become concrete tasks, and the cycle restarts on a better basis.
Frequently asked questions
How quickly does a submitted URL get indexed?
No deadline can be committed to, and any number you are quoted is a guess. Submission usually shortens the time to discovery and to the first fetch. Whether acceptance follows, and when, is decided by the search engine on the basis of the content. The log at least shows whether the fetch took place — half the answer.
Are 1,000 URLs a day enough for a mid-sized shop?
For ongoing operation, almost always. New items and changed pages stay well below that mark even in busy shops. It only gets tight with the initial intake of large inventories, with relaunches and with migrations. That is exactly when sorting in advance pays off, since the relevant share of a catalogue usually reduces to a fraction of the total.
Does the Indexing Hub replace Google Search Console?
No, and it is not meant to. Search Console remains the authority on how the search engine sees your domain. The hub adds the active part: batch submission, recursive sitemap processing and its own bot visit log. Together the two give a picture neither provides alone.
What happens if I submit the same URL repeatedly?
You spend daily allowance and gain nothing. Resubmitting an unchanged page speeds up no decision. Handing an address over again makes sense only when the content, the canonicalisation or a technical blockage has genuinely changed — after removing an accidental noindex, for instance.
How many sitemap jobs can I run at once?
Two at once, with another twenty able to sit in the queue. For a small agency with several clients that is enough in normal operation, but it requires a deliberate order. Schedule time-critical cases such as a relaunch ahead of routine maintenance, otherwise one large pass blocks the urgent one.
Conclusion: visibility starts with findability
Indexing is the underestimated bottleneck because it stays invisible when it works and does not look like a fault when it fails. No warning message, no red bar, only pages without impressions. Measure it and you find the cause; skip it and you work on copy and links while the cause sits one stage earlier.
The Indexing Hub does not change that by magic. It brings three things together: a throughput limit that forces prioritisation; an active notification method that shortens discovery time; and a log that replaces assumptions with timestamps and status codes. The limits belong to the picture — a thousand addresses a day is finite, two parallel jobs is finite, and a submission remains a submission.
That sobriety is what makes the area plannable. Someone who knows that 40,000 addresses take around 40 days at full allowance plans differently from someone who presses a button and hopes. To check how many of your pages are actually covered, open the Semalt dashboard and start with a sitemap pass; the first picture is usually more revealing than expected.