URL Normalise — strict, and it tells you what it changed
Says whether two URLs address the same thing, and names every change it made to decide. Malformed input is refused with a reason and a position instead of silently repaired: the standard parser deletes newlines, keeps a broken percent-escape as text, and encodes a raw space without saying so.
Called as web.url.normalise.
What it accepts
The URLs to normalise, and which meaning-changing normalisations you accept. Every item gets its own result: a bad URL is reported with a reason and a position, is not billed, and does not stop the ones after it.
| Field | Type | Meaning |
|---|---|---|
items | array | The URLs to normalise, in order. Results carry the index they came from. Absolute http and https only: a relative URL has no meaning without the base it is relative to, and this capability is not given one. At most 2,000 per call: at 50 ms an item that is 102 seconds, inside the platform's hard 300-second window for a synchronous call. A larger job is several calls, which is deliberate. |
alsooptional | array | Seven normalisations that cannot change which resource is addressed are always applied. These five can. Dropping www addresses a different name and it may be a different server. Sorting the query reorders parameters a server may read in order. Dropping the fragment discards what a client-side application uses to route. Pick only what you can defend. One of sorted-query, dropped-fragment, dropped-trailing-slash, dropped-www, dropped-empty-query. |
compareTooptional | string | Optional. One URL, normalised with the same settings. Every result then says whether it came out identical to it. If this one is not a URL, the run stops before charging anything: every comparison would otherwise be measured against nothing. |
What a run leaves behind
| Results, one row per item | One row per item in the order they were given, each carrying its index. A failed item is written here too, with the rule it broke and the position, so a caller can tell an item that could not be normalised from one that was never reached. |
|---|---|
| Report: what the run did, including what it did not do | delivered, succeeded, failed, notAttempted and stoppedEarly. notAttempted above zero with stoppedEarly true is a complete answer, not a truncated one: the caller's spending limit stopped the work and the run says how many it never reached. |
What it costs
- Per call
- $0.001
charged when the run opens - Per item delivered
- $0.0002
an item that fails is not billed
An item that fails is not billed, and one bad item does not stop the ones after it. What that adds up to, and what a spending limit does to a run in progress, is on the pricing page.
What would end it
The channel absorbs URL comparison as a free built-in, or a standard parser starts reporting what it repaired — at which point the difference this sells stops existing. Also the usual threshold: monthly revenue below the cost of keeping it running, two months running.
Written before it happens, on purpose. A capability that quietly stops being worth running costs its buyers more than one that says in advance how it ends.
Facts
| Version | 0.1.0 |
|---|---|
| Available | Yes, on Apify Store |
| Channel | Apify Store |
| How to call it | Console, HTTP or MCP |
| This page as Markdown | /c/url-normalise.md — the same facts without the markup, plus the full input schema and a real example of what comes back |