# URL Normalise — strict, and it tells you what it changed

> Says whether two URLs address the same thing, and names every change it made to decide. Malformed input is refused with a reason and a position instead of silently repaired: the standard parser deletes newlines, keeps a broken percent-escape as text, and encodes a raw space without saying so.

- Called as: `web.url.normalise`
- Availability: Callable on Apify Store: https://apify.com/telyvar/url-normalise
- Version: 0.1.0
- Page: https://telyvar.com/c/url-normalise/

## What it accepts

The URLs to normalise, and which meaning-changing normalisations you accept. Every item gets its own result: a bad URL is reported with a reason and a position, is not billed, and does not stop the ones after it.

| Field | Type | Required | Meaning |
| --- | --- | --- | --- |
| `items` | array (max 2,000) | yes | The URLs to normalise, in order. Results carry the index they came from. Absolute http and https only: a relative URL has no meaning without the base it is relative to, and this capability is not given one. At most 2,000 per call: at 50 ms an item that is 102 seconds, inside the platform's hard 300-second window for a synchronous call. A larger job is several calls, which is deliberate. |
| `also` | array | no | Seven normalisations that cannot change which resource is addressed are always applied. These five can. Dropping www addresses a different name and it may be a different server. Sorting the query reorders parameters a server may read in order. Dropping the fragment discards what a client-side application uses to route. Pick only what you can defend. One of: sorted-query, dropped-fragment, dropped-trailing-slash, dropped-www, dropped-empty-query. |
| `compareTo` | string | no | Optional. One URL, normalised with the same settings. Every result then says whether it came out identical to it. If this one is not a URL, the run stops before charging anything: every comparison would otherwise be measured against nothing. |

## Example input

```json
{
  "items": [
    "HTTPS://EXAMPLE.com:443/a/./b/../c?b=2&a=1#top",
    "https://example.com/a/c?b=2&a=1#top",
    "https://example.com/%zz"
  ],
  "also": [],
  "compareTo": "https://example.com/a/c?b=2&a=1#top"
}
```

## What comes back

One row per item, in the order they were given, each carrying the index it came from. A row for an item that failed is written too, with the rule it broke and where.

| Field | Type | Always present |
| --- | --- | --- |
| `index` | integer | yes |
| `ok` | boolean | yes |
| `url` | string | no |
| `scheme` | string | no |
| `host` | string | no |
| `hostAscii` | string | no |
| `changed` | string | no |
| `sameAsCompare` | boolean | no |
| `reason` | string | no |
| `detail` | string | no |

## Example output

What this capability really returns for the example input above. It is not written by hand: the capability's own test re-runs it and fails if this disagrees.

```json
{
  "results": [
    {
      "index": 0,
      "ok": true,
      "url": "https://example.com/a/c?b=2&a=1#top",
      "scheme": "https",
      "host": "example.com",
      "hostAscii": "example.com",
      "changed": [
        "lowercased-scheme",
        "lowercased-host",
        "dropped-default-port",
        "resolved-dot-segments"
      ],
      "sameAsCompare": true
    },
    {
      "index": 1,
      "ok": true,
      "url": "https://example.com/a/c?b=2&a=1#top",
      "scheme": "https",
      "host": "example.com",
      "hostAscii": "example.com",
      "changed": [],
      "sameAsCompare": true
    },
    {
      "index": 2,
      "ok": false,
      "reason": "bad-percent-encoding",
      "detail": "position 20 is a % not followed by two hexadecimal digits, which the standard parser keeps as literal text"
    }
  ],
  "report": {
    "delivered": 3,
    "succeeded": 2,
    "failed": 1,
    "notAttempted": 0,
    "stoppedEarly": false
  }
}
```

## What it costs

- $0.001 per call, charged when the run opens
- $0.0002 per item delivered
- An item that fails is not billed, and one bad item does not stop the ones after it
- A call of 100 items therefore costs $0.021

## What would end it

The channel absorbs URL comparison as a free built-in, or a standard parser starts reporting what it repaired — at which point the difference this sells stops existing. Also the usual threshold: monthly revenue below the cost of keeping it running, two months running.

Written before it happens, on purpose.

---

Telyvar — https://telyvar.com/ · How to call one: https://telyvar.com/docs.md · Prices: https://telyvar.com/pricing.md
