emjupy design

Architecture

emjupy speaks to a Jupyter server the way the browser does: HTTP for files and kernels, WebSockets for a kernel's messages and the language server's. Everything goes over the server's one port, so a remote server needs one SSH tunnel and nothing started on the far side.

 Emacs, on this machine                         Jupyter server, here or remote
┌─────────────────────────────────────┐        ┌───────────────────────────────────────┐
│ notebook buffer       (emjupy-mode) │  HTTP  │ Contents API ──── .ipynb files        │
│   cells ⇄ structs ⇄ .ipynb JSON ────┼───────▶│ Sessions and Kernels APIs             │
│     │ outputs, widget controls      │        │                                       │
│     ▼                               │   WS   │                          ZMQ          │
│   kernel connection ────────────────┼───────▶│ /api/kernels/…/channels ──▶ kernel    │
│                                     │   WS   │                                       │
│ shadow buffer (.py) ◀── Eglot ──────┼───────▶│ /lsp/ws/… ── jupyter-lsp ──▶ pylsp    │
│                                     │        │                                       │
│ figure window ◀── page ◀── bridge   │        └───────────────────────────────────────┘
└─────────────────────────────────────┘
   the bridge relays a widget page's messages over the
   kernel connection above
  • The notebook buffer is the notebook: cells drawn as text, kept in step with the cell structs, which are what is saved as .ipynb.
  • The kernel connection carries execution and its output; widgets' messages too.
  • The shadow buffer is every code cell as one Python file, for Eglot, whose language server runs beside the kernel through jupyter-lsp.
  • The figure window shows what a text buffer cannot – plotly figures, widgets that draw themselves – and a widget's page talks back through the bridge, a WebSocket server in Emacs, not to the Jupyter server.

The shadow buffer

A notebook is not valid Python, but a JSON document holding cells. There are n Python fragments with prose between them, and a name defined in, say, cell 3 is used in cell 7. A language server handed one cell at a time would see seven unrelated files. So instead, emjupy keeps a shadow buffer: one ordinary Python file holding every code cell concatenated, each preceded by a marker comment naming its cell. eglot attaches to that file and sees exactly what it expects – one module, in order, with all the definitions in it.

The shadow buffer is hidden: it exists so that a language server has a well-formed document to read. C-c ' shows it, for editing several cells at once; C-c C-c there writes the changes back into the cells.

Everything that needs the server is delegated there at the matching position:

  • completion and eldoc, through completion-at-point-functions and eldoc-documentation-functions;
  • M-. and friends, through an xref backend that forwards to eglot's own and maps the answers back onto cells;
  • eglot-rename and the other eglot commands, advised to run in the shadow buffer and pull their edits back.

Because the shadow buffer holds every cell at once, this is not merely a workaround: a rename fixes references in all cells, and M-. on a call in one cell lands on the definition in another.

The language server beside the kernel

The language server that answers must run where the kernel runs: only there can it see your own modules and the Python the code runs against. emjupy talks to it the way JupyterLab does, through jupyter-lsp on the Jupyter server, over a WebSocket on the same connection the notebook uses – so it works through the same SSH tunnel, with nothing started on the remote by hand.

The shadow file itself stays on this machine – beside the notebook when the notebook is on this machine too, in a temporary directory when it is not – and nothing on the far side reads it. Its contents go over the socket as the LSP document, and every path in the traffic is rewritten in flight – the shadow's local directory to the kernel's working directory on the way out, and back on the way in – so the server resolves import mylib beside the notebook, where the module is. The server is also told that directory explicitly, as the place to look for modules.

When that server cannot be reached, emjupy falls back to one on this machine; what that costs is in Language server support.

Mapping positions

Point in a cell maps to a position in the shadow buffer by finding that cell's marker and adding the offset within the cell; results coming back map the other way. Everything the server returns is in shadow-file coordinates, so xref results inside the shadow are rewritten into positions in the notebook.

A location anywhere else is a file on the kernel's machine, when the server beside the kernel answered, and is named there over TRAMP: the host comes from emjupy-remote-root or from the SSH tunnel the server is reached through. When the answer came from the local fallback, its paths already name this machine's files and are left alone.

Output

Output arrives as messages on the kernel's WebSocket, each naming the request it belongs to. A request is finished when two things have come: the reply, on the shell channel, and the status message that follows the request's last output, on the iopub channel. The protocol does not order one channel against the other, so an output can arrive after its reply; waiting for both is what keeps it.

Stream text continues the output it belongs to rather than starting a new one. Consecutive text on one stream is one output, and text on a stream whose line is still open – no newline yet – continues that line even if the other stream wrote in between. That is a terminal's rule, and the one a progress bar relies on: each update begins with a carriage return, and carriage returns are applied as a terminal would.

How wide output is drawn, and when the kernel is told the width, is in Using a notebook.

Output is redrawn at most a few times a second, however many messages arrive, and only in the cell it belongs to.

Closing a notebook

Killing a notebook's buffer closes what it opened: the connection to its kernel, its language-server connection and its shadow buffer. The kernel itself keeps running on the server – it outlives the buffer by design, and the next login adopts it. A language server another notebook still uses is left running.

The connection is ended directly: the shadow buffer is killed first, which is how Eglot closes the document, then the server is marked as shut down on purpose – or Eglot would start a replacement – its pending requests are dropped, and it is taken out of Eglot's table at once, so a notebook opened a moment later is not handed a dying server. None of it waits: closing takes no measurable time.

Redrawing

Most changes redraw only the part of the buffer they affect: output arriving, running a cell (which clears its previous output first), inserting, deleting, splitting, merging and moving cells, hiding and clearing output. The rest of the buffer moves by a known amount, so anything holding a position there – markers, other windows, the fake cursors of multiple-cursors – stays where it was.

A few redraw the whole notebook: changing a cell's type, emjupy-re-render, rendering a markdown cell, and undoing a structural change. They erase the buffer and draw it again from the cells, which moves every held position to the start.

Undo

Every cell command – inserting, deleting, moving, splitting, joining, changing a cell's type, clearing or hiding output, yanking a cell – is one undo step, and redo takes it forward again. They all work the same way: the cells are snapshotted before the command changes anything, and undoing the step restores the snapshot. Restoring records the state it replaced, so redo is the same operation the other way.

Three things make that sound. The snapshot is taken first: a command that changed a cell and then redrew left the redraw to record the state after the change, so undoing it restored the change. A snapshot is written back into the same cell objects, not copies, since the notebook and every other undo entry refer to them. And while a cell command runs, no earlier undo entry is shifted or dropped: undoing the step restores exactly the text those entries describe, and they are only replayed after it. Shifting them instead could not be done right – undoing a deletion reinserts the cell, and the shift that makes room for it moved the cell's own entries into the next one.

Undo entries hold plain buffer positions, so a redraw that is not an undo step – output arriving, say – keeps them meaningful directly: entries before the change are untouched, those after it are shifted by the amount it moved things, and those inside the redrawn region are dropped. It changes the entries in place, since an undo already under way replays from a list that is shared, not copied.

Use the browser's transport, not ZMQ

When you run jupyter server --no-browser --port=8888, the server exposes:

  • GET /api/kernels → list running kernels
  • POST /api/kernels → start a kernel
  • GET /api/sessions → list sessions (kernel ↔ notebook pairs)
  • POST /api/sessions → create or attach a session
  • GET /api/contents → list notebooks
  • GET /api/contents/<path> → fetch notebook JSON
  • PUT /api/contents/<path> → save notebook JSON
  • WS /api/kernels/<id>/channels → the Jupyter protocol over WebSocket
  • WS /lsp/ws/<server> → a language server, through jupyter-lsp

This is what every browser-based client (JupyterLab, Classic Notebook) uses. It works transparently through an ssh tunnel, survives reconnects, and needs only url.el (built in) and websocket.el.