Kjeks
← All docs

Kjeks Scanner

A standalone Node + Playwright scanner. Crawls your pages under every consent state and reports the trackers that actually fire.

Standalone Node 20+ CLIUses PlaywrightFeeds the Kjeks core inventory Source on GitHub →

101 First scan

The scanner runs outside WordPress. It drives a real Chromium browser through each consent state and records the network requests, cookies, and storage that appear — catching trackers injected by themes, embeds, or other plugins that are easy to miss by hand.

Prerequisite: the Kjeks core plugin must be installed and active on the site you scan. The scanner reproduces each consent state through Kjeks' own kjeks_consent cookie and storage, and imports results back through the core plugin's scan-config and import surfaces — without it there is nothing to configure, block, or import against.
# Fastest — run straight from npm, no clone (Node 20+)
npx kjeks-scanner --url https://example.com/ --blog-id 1 --out scan

# Or install it globally, then use the kjeks-scanner / kjeks-scan command
npm install -g kjeks-scanner
kjeks-scanner --url https://example.com/ --blog-id 1 --out scan

# From a clone (for development)
npm install
npx playwright install chromium
node src/cli.js --url https://example.com/ --blog-id 1 --out scan
The scanner is published on npm as kjeks-scanner, so npx kjeks-scanner needs no clone. The first run downloads a pinned Chromium (~100 MB) via Playwright; later runs reuse it. After a global install, both kjeks-scanner and kjeks-scan work.

The six consent states

Every path is loaded once per state, each in a fresh browser context:

StateConsent injected
before-choiceNone — the banner hasn't been answered.
reject-allEverything optional denied.
only-preferencesOnly preferences granted.
only-analyticsOnly analytics granted.
only-marketingOnly marketing granted.
accept-allAll optional categories granted.
For states with a choice, the scanner writes the kjeks_consent cookie and localStorage value before page scripts run — so blocking is evaluated exactly as a returning visitor would experience it.

201 Config & authentication

Config file

For anything beyond a single URL, describe your sites in a config file and pass --config. Each site lists the paths to crawl and optional scenarios (scripted click/wait steps for gated content like video embeds):

{
  "sites": [
    {
      "url": "https://example.com/",
      "blog_id": 1,
      "policy_version": 1,
      "paths": [ "/", "/about" ],
      "scenarios": [
        {
          "name": "open a gated video",
          "steps": [
            { "action": "click", "selector": ".kjeks-embed__load" },
            { "action": "wait", "ms": 1000 }
          ]
        }
      ]
    }
  ]
}

Generate it from WordPress

Don't hand-write the site list — the core plugin can emit it. Use the CLI, or pull it from the REST route with credentials:

# Generate the config from WordPress, then scan it
wp kjeks scan-config --output=config.json
npx kjeks-scanner --config config.json --out scan

# Pull config straight from a live site (shared scanner key)
KJEKS_SCAN_KEY='<key from: wp kjeks scan-key --generate>' \
  npx kjeks-scanner --config-url https://example.com/wp-json/kjeks/v1/scan-config --out scan

Invocation forms

FlagPurpose
--url + --blog-idScan one URL without a config file.
--config <file>Scan sites from a local config file.
--config-url <url>Fetch the config from a live site (needs KJEKS_USER + KJEKS_APP_PASSWORD).
--overlay <file>Merge extra paths/scenarios by blog_id.
--concurrency <n>Sites scanned in parallel (default 3).
--per-host <n>Parallel scans sharing one hostname (default 2) — politeness for subdirectory multisites.
--fullScan the server selection as-is; skip re-scanning pages that previously produced a tracker.
--import [<url>]After scanning, POST observations to the import endpoint in the same run.
--out <dir>Output directory (default scan).
--endpoint <CDP>Connect to an existing browser over the DevTools protocol.

301 Output & CI

What it writes

One stable JSON file per site is written to <out>/<host>[_<path>].json. Each file records every state plus a flattened list of observations:

{
  "host": "example.com",
  "url": "https://example.com/",
  "blog_id": 1,
  "states": {
    "before-choice": { "cookies": [], "scripts": [], "iframes": [], /* … */ },
    "reject-all":     { /* … */ },
    "accept-all":     { /* … */ }
  },
  "observations": [
    { "name": "_ga", "storage_type": "cookie", "party": "third",
      "domain": ".example.com", "retention": "…",
      "triggered_by": ["accept-all", "only-analytics"],
      "source_urls": ["/", "/blog/hello-world/"] }
  ]
}

Per state, the scanner captures:

cookies, localStorage, sessionStorage, indexedDB, thirdPartyHosts, beacons, scripts, iframes, and redirects. Each observation also records triggered_by — the consent state that caused it — and source_urls, the page(s) it actually loaded on, so you can see exactly which choice and which page produced a tracker. (The kjeks_consent cookie itself is excluded from results.)

Targeted re-scans. The core plugin auto-selects representative URLs per site (home, newest post/page, embed-bearing pages, capped). On each run the scanner also re-scans every page that previously produced a tracker (from source_urls), so sampling never drops a known-tracker page. Pass --full to scan the server selection as-is.

Diffing & exit codes

On each run the scanner compares the new file with the previous one and exits with status 1 when anything changed. That makes it a natural CI gate: a non-zero exit means a new tracker appeared and your inventory needs review.

Run it in GitHub Actions

The easiest way to run the scanner on a schedule: drop this workflow into .github/workflows/kjeks-scan.yml in any repo you already own. It uses npx (no clone, no npm ci, no lockfile), scans, imports the observations, and uploads the results as a downloadable artifact.

name: Kjeks discovery scan

on:
  schedule:
    - cron: '0 3 * * 1' # Weekly, Monday 03:00 UTC
  workflow_dispatch:

permissions:
  contents: read

jobs:
  scan:
    runs-on: ubuntu-latest
    # A full multisite scan takes ~2h; cap the job well under GitHub's 6h
    # ceiling so a stalled run can never burn a whole runner-day.
    timeout-minutes: 180
    steps:
      - uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
        with:
          node-version: 20
      - run: npx playwright install --with-deps chromium
      - name: Run discovery scan and import
        id: scan
        # Slightly under the job cap so the artifact upload still runs on a hang.
        timeout-minutes: 170
        # Non-zero exit = a subsite changed / a site errored / observations were
        # imported (the normal outcome). Keep the job green and still upload below.
        continue-on-error: true
        env:
          KJEKS_SCAN_KEY: ${{ secrets.KJEKS_SCAN_KEY }}
          # Kept in env (not expanded inline in run:) to avoid a script-injection surface.
          KJEKS_SITE_URL: ${{ secrets.KJEKS_SITE_URL }}
        run: >-
          npx kjeks-scanner
          --config-url "$KJEKS_SITE_URL/wp-json/kjeks/v1/scan-config"
          --import --out scan
      - uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
        with:
          name: kjeks-scan
          path: scan/

Add two repository secrets — KJEKS_SITE_URL and KJEKS_SCAN_KEY (generate it with wp kjeks scan-key --generate) — and you're done. The key travels in the X-Kjeks-Key header, which survives proxies/CDNs that strip Authorization (a common cause of 401 rest_not_logged_in). Basic auth with KJEKS_USER + KJEKS_APP_PASSWORD still works as a fallback. Never commit the key.

Why continue-on-error? The scanner exits non-zero when a subsite changed, a site errored, or observations were imported — the normal outcome of a successful run that found something. Without continue-on-error: true on the scan step, that expected exit code fails the job and skips the artifact upload. Keep it green and review the imported observations in Network Admin → Cookie Consent; drop continue-on-error only if you want CI to go red whenever the scan detects a change.
Want the scan to diff against a committed baseline and flag regressions? That variant needs a checkout to commit updates back. See the scanner README for the full baseline-tracking workflow and all the details.
Why the timeout-minutes? A full multisite scan takes roughly two hours, so the workflow caps the job at 180 and the scan step at 170 — safely above a normal run but well under GitHub's 6-hour ceiling, so a stalled run fails fast instead of burning a whole runner-day. The scanner itself (0.3.6+) already bounds every network wait — config fetch, import POST, cookie read, and page evaluation all time out — so a slow or unreachable site is skipped rather than hanging the run; the job timeout is just a belt-and-suspenders backstop.
Behind a login gate? If the site runs Restricted Site Access (or similar), it redirects anonymous visitors — including the REST API — to wp-login.php, so the scanner gets an HTML login page instead of JSON. Kjeks 1.0.7+ lets any request carrying a valid KJEKS_SCAN_KEY through, and the scanner sends that key on every request (REST and page loads), so it can scan the restricted site while the public stays blocked. The key alone grants that access — rotate or clear it from Settings when you're done.

Close the loop

Once you've reviewed the results (optionally with the AI Reviewer), import them back into WordPress — either as a separate step, or in the same run with --import:

# Feed reviewed observations back into the inventory
wp kjeks import scan/example.com.json --blog_id=1
The scanner has no WordPress hooks or options — it integrates purely through the core plugin's scan-config and import surfaces. See the core 301 guide for those.