# The walk a faster kernel has to reproduce

Anything that attacks a rung of this ladder, in CUDA or anything else, has to walk exactly this walk,
or two implementations will never collide with each other and the work is wasted. The reference is
`../rho.mjs`, and it is the only definition that counts; this file is the same thing in words.

## The group

A rung gives a prime `p`, a curve `y^2 = x^3 + a x + b` over `F_p`, a prime order `n`, and two points
`P` and `Q`. Everything is derived from the rung's `seed` by `../../app/ec.mjs`, and every number is
in `data/ladder.json`. A point is held canonically: of `(x, y)` and `(x, p - y)` the walk always
keeps the one with `y <= p / 2`, which is the negation map and halves the space.

## The step table

Thirty two steps, `M[j] = u_j * P + v_j * Q`, with

    u_j = SHA-256 counter mode over "<seed>/step/<j>/u", reduced mod n
    v_j = SHA-256 counter mode over "<seed>/step/<j>/v", reduced mod n

The counter mode is the one in `hashInt`: concatenate `SHA-256("<label>|0")`, `SHA-256("<label>|1")`,
... until there are 32 bytes, take the first 32, read them big endian.

## The walk

From a point `X = a*P + b*Q`:

    j   = low five bits of X.x
    X'  = X + M[j],  a' = a + u_j mod n,  b' = b + v_j mod n
    if X'.y > p / 2:  X' = -X',  a' = -a' mod n,  b' = -b' mod n

The negation map makes short fruitless cycles. Keep the last sixteen points of each walk; when the
new point is already in that trail, the walk is cycling. Escape by doubling the point of the cycle
with the SMALLEST x, coefficients doubled with it. The escape has to be canonical like this, or two
walks that merged and then met the same cycle would escape to different places and the merge would
be thrown away.

## A distinguished point

`X.x` with its low `d` bits zero, where `d` is chosen so a card reports a few thousand points an
hour. Report `(x, a, b)`. Two reports of the same `x` with different `b` give

    k = (a1 - a2) / (b2 - b1) mod n

and `k * P == Q` is the only check that matters.

## Reporting

POST to the url the keeper gives you, as often as `--report-every` says:

```json
{ "bits": 69, "seed": "ECDLP-69/69/115", "ops": 81234567, "dps": 1204,
  "answer": null, "took": "2h 14m", "opsPerSecond": 10123456, "final": false }
```

`answer` is the decimal `k` once two walks have met, and `final` is true on the last report.

## What ships today

`../worker.mjs` is the worker the keeper dispatches: the reference walk across every core of the
rented machine. It is correct and it is in the test suite. A CUDA kernel is the first speed step and
the first thing the research agents are pointed at; it is not written yet, because there is no CUDA
toolchain and no card on the machine this project was built on, and an unverified kernel is worth
less than none. Build one against this file and the reference will agree with it.
