emjupy design
Architecture
emjupy speaks to a Jupyter server the way the browser does: HTTP for files and kernels, WebSockets for a kernel's messages and the language server's. Everything goes over the server's one port, so a remote server needs one SSH tunnel and nothing started on the far side.
Emacs, on this machine Jupyter server, here or remote ┌─────────────────────────────────────┐ ┌───────────────────────────────────────┐ │ notebook buffer (emjupy-mode) │ HTTP │ Contents API ──── .ipynb files │ │ cells ⇄ structs ⇄ .ipynb JSON ────┼───────▶│ Sessions and Kernels APIs │ │ │ outputs, widget controls │ │ │ │ ▼ │ WS │ ZMQ │ │ kernel connection ────────────────┼───────▶│ /api/kernels/…/channels ──▶ kernel │ │ │ WS │ │ │ shadow buffer (.py) ◀── Eglot ──────┼───────▶│ /lsp/ws/… ── jupyter-lsp ──▶ pylsp │ │ │ │ │ │ figure window ◀── page ◀── bridge │ └───────────────────────────────────────┘ └─────────────────────────────────────┘ the bridge relays a widget page's messages over the kernel connection above
- The notebook buffer is the notebook: cells drawn as text, kept in
step with the cell structs, which are what is saved as
.ipynb. - The kernel connection carries execution and its output; widgets' messages too.
- The shadow buffer is every code cell as one Python file, for Eglot,
whose language server runs beside the kernel through
jupyter-lsp. - The figure window shows what a text buffer cannot – plotly figures, widgets that draw themselves – and a widget's page talks back through the bridge, a WebSocket server in Emacs, not to the Jupyter server.
The shadow buffer
A notebook is not valid Python, but a JSON document holding cells.
There are n Python fragments with prose between them, and a name
defined in, say, cell 3 is used in cell 7. A language server handed
one cell at a time would see seven unrelated files. So instead,
emjupy keeps a shadow buffer: one ordinary Python file holding
every code cell concatenated, each preceded by a marker comment naming
its cell. eglot attaches to that file and sees exactly what it
expects – one module, in order, with all the definitions in it.
The shadow buffer is hidden: it exists so that a language server has a
well-formed document to read. C-c ' shows it, for editing several
cells at once; C-c C-c there writes the changes back into the cells.
Everything that needs the server is delegated there at the matching position:
- completion and eldoc, through
completion-at-point-functionsandeldoc-documentation-functions; M-.and friends, through an xref backend that forwards toeglot's own and maps the answers back onto cells;eglot-renameand the othereglotcommands, advised to run in the shadow buffer and pull their edits back.
Because the shadow buffer holds every cell at once, this is not merely
a workaround: a rename fixes references in all cells, and M-. on a
call in one cell lands on the definition in another.
The language server beside the kernel
The language server that answers must run where the kernel runs: only
there can it see your own modules and the Python the code runs against.
emjupy talks to it the way JupyterLab does, through jupyter-lsp on
the Jupyter server, over a WebSocket on the same connection the notebook
uses – so it works through the same SSH tunnel, with nothing started on
the remote by hand.
The shadow file itself stays on this machine – beside the notebook when
the notebook is on this machine too, in a temporary directory when it is
not – and nothing on the far side reads it. Its contents go over the socket as
the LSP document, and every path in the traffic is rewritten in flight
– the shadow's local directory to the kernel's working directory on the
way out, and back on the way in – so the server resolves import mylib
beside the notebook, where the module is. The server is also told that
directory explicitly, as the place to look for modules.
When that server cannot be reached, emjupy falls back to one on this
machine; what that costs is in Language server support.
Mapping positions
Point in a cell maps to a position in the shadow buffer by finding that
cell's marker and adding the offset within the cell; results coming
back map the other way. Everything the server returns is in
shadow-file coordinates, so xref results inside the shadow are
rewritten into positions in the notebook.
A location anywhere else is a file on the kernel's machine, when the
server beside the kernel answered, and is named there over TRAMP: the
host comes from emjupy-remote-root or from the SSH tunnel the server
is reached through. When the answer came from the local fallback, its
paths already name this machine's files and are left alone.
Output
Output arrives as messages on the kernel's WebSocket, each naming the request it belongs to. A request is finished when two things have come: the reply, on the shell channel, and the status message that follows the request's last output, on the iopub channel. The protocol does not order one channel against the other, so an output can arrive after its reply; waiting for both is what keeps it.
Stream text continues the output it belongs to rather than starting a new one. Consecutive text on one stream is one output, and text on a stream whose line is still open – no newline yet – continues that line even if the other stream wrote in between. That is a terminal's rule, and the one a progress bar relies on: each update begins with a carriage return, and carriage returns are applied as a terminal would.
How wide output is drawn, and when the kernel is told the width, is in Using a notebook.
Output is redrawn at most a few times a second, however many messages arrive, and only in the cell it belongs to.
Closing a notebook
Killing a notebook's buffer closes what it opened: the connection to its kernel, its language-server connection and its shadow buffer. The kernel itself keeps running on the server – it outlives the buffer by design, and the next login adopts it. A language server another notebook still uses is left running.
The connection is ended directly: the shadow buffer is killed first, which is how Eglot closes the document, then the server is marked as shut down on purpose – or Eglot would start a replacement – its pending requests are dropped, and it is taken out of Eglot's table at once, so a notebook opened a moment later is not handed a dying server. None of it waits: closing takes no measurable time.
Redrawing
Most changes redraw only the part of the buffer they affect: output
arriving, running a cell (which clears its previous output first),
inserting, deleting, splitting, merging and moving cells, hiding and
clearing output. The rest of the buffer moves by a known amount, so
anything holding a position there – markers, other windows, the fake
cursors of multiple-cursors – stays where it was.
A few redraw the whole notebook: changing a cell's type,
emjupy-re-render, rendering a markdown cell, and undoing a structural
change. They erase the buffer and draw it again from the cells, which
moves every held position to the start.
Undo
Every cell command – inserting, deleting, moving, splitting, joining, changing a cell's type, clearing or hiding output, yanking a cell – is one undo step, and redo takes it forward again. They all work the same way: the cells are snapshotted before the command changes anything, and undoing the step restores the snapshot. Restoring records the state it replaced, so redo is the same operation the other way.
Three things make that sound. The snapshot is taken first: a command that changed a cell and then redrew left the redraw to record the state after the change, so undoing it restored the change. A snapshot is written back into the same cell objects, not copies, since the notebook and every other undo entry refer to them. And while a cell command runs, no earlier undo entry is shifted or dropped: undoing the step restores exactly the text those entries describe, and they are only replayed after it. Shifting them instead could not be done right – undoing a deletion reinserts the cell, and the shift that makes room for it moved the cell's own entries into the next one.
Undo entries hold plain buffer positions, so a redraw that is not an undo step – output arriving, say – keeps them meaningful directly: entries before the change are untouched, those after it are shifted by the amount it moved things, and those inside the redrawn region are dropped. It changes the entries in place, since an undo already under way replays from a list that is shared, not copied.
Use the browser's transport, not ZMQ
When you run jupyter server --no-browser --port=8888, the server
exposes:
GET /api/kernels→ list running kernelsPOST /api/kernels→ start a kernelGET /api/sessions→ list sessions (kernel ↔ notebook pairs)POST /api/sessions→ create or attach a sessionGET /api/contents→ list notebooksGET /api/contents/<path>→ fetch notebook JSONPUT /api/contents/<path>→ save notebook JSONWS /api/kernels/<id>/channels→ the Jupyter protocol over WebSocketWS /lsp/ws/<server>→ a language server, throughjupyter-lsp
This is what every browser-based client (JupyterLab, Classic Notebook)
uses. It works transparently through an ssh tunnel, survives
reconnects, and needs only url.el (built in) and websocket.el.
