Direct Editing: what it took to let people edit a running prototype
AI Systems / Performance / Protocols
- Rendering 1000+ iframes inside a canvas all of which have their own heap and event loop running, not possible if done naively
- Finding a solution where flipping thumbnails and mounted frames are instant
- 1,001 screens rendered maintaining full functionality, at 60 fps on an M1 Air. A user selects an element in any screen, edits it live, and sends the whole intent to the agent as one request
Figr is a design agent that lives on your product. It plans, writes documents and builds working prototypes on a canvas that a team works on at the same time. So naturally our way of showing these interactive webapps was an iframe on the screen. There is a technical limitation we all ignored for long though. A handful of iframes was fine. A hundred was a browser tab at 7 GiB memory consuption and an “Aw, Snap”.
Direct editing was a feature that needed us right up against that limit, and then some on top for the editing overhead itself. The feature sounded simple; open a prototype and see every state as a screen laid out like a Figma file; panning and zoom; hover any element in any screen; click it, change its padding or its text; see the change live, and then hand the whole set of changes to the agent as one request. It is live on our product today; this is the story of how.
1. what a running screen costs
The first question you may have is why a hundred iframes is a problem at all. Well, iframes load other sites embedded in another webapp, and loading any random website is not a very safe thing, so chromium isolates all iframes into a separate sandbox based on certain conditions. Our case however was not hurt by this because chromium isolates by site, not by origin. So that was my first gotcha, if we are not paying the sandbox cost, why are screens chugging so much memory. chromium was already sharing one single thread for rendering everything. I found though, what is not shared is the heap (the memory the webapp actually uses); each document gets its own realm and evaluates the whole prototype bundle again, so every object the app creates on boot exists once per frame. 6 MiB per frame on macOS and around 25 MiB on archlinux, where a cgroup (a linux fence around how much memory a process may use) count decides how much exactly. Nothing frees that memory except for an unmount.
I was not going to accept that. I searched hard for a way out. I was searching for a way to share the heap; looked at using shadow DOM instead of iframes; however it will fall short in what we wanted; media queries stay page-level, portals escape the frame bounds and focus behaviour is quirky. To make it work, the source code required extra wiring and a special router, I did not want to add shackles to our agent. Heck, I was even considering using WebAssembly. However, ultimately I realised none of them will work optimally; so I decided to think from a UX point of view and then engineer solutions. I set a target for myself: 1000 screens at 60 FPS on a 4 GiB box and a benchmark that will run on every iteration.
2. the browser lied about its own memory, twice
So only a few frames can be live at any given time, but the user should never feel that; everything on the canvas has to look rendered. That means something has to decide which frames get to live, and that decision needs a budget. The obvious input is navigator.deviceMemory; the browser tells you how much RAM the machine has, you derive a count from it right? Except, on my linux box capped at 4 GiB through a cgroup, it reported 16. The budget resolved to 96 live frames, the board happily mounted 108 of them, and over the next few minutes I watched the browser hit it’s memory ceiling 688 times before the kernel killed it at 6m56s. Forcing the value to 4 held the exact same board at 2,391 MiB with zero memory reclaim hits. So the mechanism was right; however, it was just lied to; the API reports the physical memory of the host, a container limit is simply invisible to it. I could have accepted this since almost no one uses containers for browsers, except, now you have agents who run their own browsers in virtual hardware provided by many. So I decided to make it resilient and not rely on a half baked measure.
The budget became a closed-loop of defensive checks. The memory limit hint is only a start; then, if three frames in a row go live and never report ready within 15 seconds, the budget halves; because a starved or killed renderer is the only memory-pressure signal that I can get in the abstracted world that my javascript lives in. Sustained success nudges the budget upwards.
Then the second lie. chromium’s own memory-pressure evaluator reads the host’s /proc/meminfo, so inside a container it never feels the cap, and on desktop linux there is no evaluator at all. Garbage accumulates until the kernel kills the browser. A memory dump of the 1001-screen board showed 141 MiB of live objects and up to 1000 MiB of garbage; nearly all of it is churn, from how nodes get painted. I do not have a knob to trigger garbage collection. I had to stop producing garbage: a demoted screen node became one element with one transform, and the garbage fell from 28449 to 5978 (paint-property nodes).
3. a picture of every screen
For frames that are not promoted to be live based on the budget, I’ve got to show at least something, it can’t be blank. My first iteration on this was to capture the base64 blob after mounting the frame. This meant though, every tile/screen must boot, paint, then paint onto a canvas, then get demoted back if the budget said so. That was expensive to build, and it was a bad UX for the viewer. The photograph had to come from somewhere that did not depend on what any browser rendered.
So every build of a prototype is rendered asynchronously after it is ready; every screen at 3 resolutions for zoom, into three WebP sizes, by a fleet of headless browsers behind a queue. The queue is not a queue in the usual sense. We only prefer the latest versions of a prototype since previous versions are seen only while peeking an older version history. The newest job is prioritised first. The first real 1000-screen job that I ran killed its browser at 55s and captured 179 screens; so I optimized the worker configuration to relaunch on crashes, divide memory properly and re-run only what it lost. I was able to achieve 60s render-then-capture-then-save-to-CDN time for all the 1000 screens at an acceptable cost for us.
4. never a blank pixel
Now I had pictures and live frames sitting next to each other on the same canvas, and the swap between the two is where the eye catches everything; a white flash, a blurry tile that snaps a second later, a frame that jumps a few pixels the moment it goes live. Every iteration I ended up making came out of one of those defects; so I will enumerate them as they happened:
The small thumbnail is always painted underneath, no matter what. A tile can be blurry, it can never be white, that is the contract. Then nothing is swapped before it is ready to be seen; a live frame is never replaced by its picture until the picture has actually decoded (img.decode() is the whole trick); a picture is not replaced by a live frame until that frame has painted a first frame, which is not the same thing as the prototype saying it is ready; during testing I learned that the handshake fires well before the first real paint, when I saw white flicker sitting there for a moment before frames are promoted.
The crisp image is the browser’s job, not mine. I started writing the algorithm and spent a day hacking around; before deciding that blurry beats exact; I reverted all of it and went back to plain srcset. I did have to wire up the size detection though because we are a canvas; and zooms are done using CSS transforms which browsers don’t take into account during src selection. For the interested: naturalWidth on a srcset image is density corrected and will confidently tell you that a 750px file is 105px wide, which cost me an entire afternoon believing that the worker was broken.
5. keeping it smooth
Frames became cheap now; you would think that’s all. I went on to the M1 machine and expected buttery smoothness. It was not. I did a profile and the results said that 1.3% of the thread was panning and everything else was chromium painting and invalidating, so there was nothing left to win in react and I had to find out what I was asking the browser to repaint.
The biggest one surprised me. Our node overlays (name tags, section borders, hover and selection outlines) scaled inversely with the zoom through one CSS custom property written on the canvas root; that single write cost us a lot on every single pan event; even if nothing read the chrome, it cost 35.8ms of time per event. Writing the value was the cost; so I had to refactor the defaults of the xyflow canvas we were using and made one single overlay that reflected the whole write instead of on a per-node basis. So long tasks while idle went from 75 to 1 and more importantly tasks during panning went from 8 to 0. The canvas library had its surprises waiting. Its viewport hook could hold only one listener at a time, so every live frame that mounted stole it from my scheduler and pictures quietly stopped becoming iframes. One event bus per canvas fixed it. There were more issues in how it handled fitView and provided information about gesture start-end events which I had to work around.
The end result was a 1001-screen board on an 8 GiB M1 Air; opening went from 63 live frames to 3, settled went from 12.8 to 60 FPS, and selecting the last screen went from 14s to 122ms.
6. the screen you touched stays alive
A picture preserves what a screen looked like, it cannot resume the document. If you interacted with a screen, typed into it, navigated inside it, or you have edits pending in it, that iframe keeps the very same document through deselection, panning away and zooming out. Nothing demotes a frame until the reason keeping it alive is released, and interactions never release.
Another requirement was carrying interactive states across view modes, switching between the canvas and the full prototype view, which unmounts the canvas entirely. The solution was to keep exactly one iframe element per prototype, owned outside both views, and it is physically moved between them with Element.moveBefore(), it relocates a node without the remove-and-insert that resets an iframe. State, route and scroll survive the move; browser support is not strong though so for others the frame remounts.
7. editing through the picture
Everything above is what it took to just show the prototype efficiently. Editing live nodes in the prototypes asked a lot more. You want to hover over any single tile and would expect the selection box to appear and clicking should select that very element. But here’s the catch; a picture has no DOM to ask for all this. I knew this was a requirement and so I had worked the place for this during the rasterization saga of the canvas; the headless browser would pre-compute the entire DOM geometry of each screen and store that as data. When the canvas loads, there is an emulation mechanism where the frame pretends to be a live prototype and speaks about where each box should be during hover and selection states. When a box is selected, the canvas brings that tile to life instantly so you never notice that you had selected a picture and not an iframe.
Editing is where things get a lot tough though. Each iframe that we mounted was cross-origin, and I cannot read their DOM and I cannot script them. They are their own react webapps that re-render, remount and get rebuilt underneath whatever the agent decides, and the same component renders in many screens at once. Mainly three ideas carry the whole design for this; for all of them we had to inject a harness script at build time that would respond and receive our commands.
Identity of each element is stamped at build time. Every element gets a marker derived from where it sits in the source code, the same marker the geometry was keyed on, and every component forwards the scope of itself onto child elements, so nothing walks the DOM at selection time. The markers are idempotent so they survive rebuilds if their input is the same. This way, we have a reliable way to exactly target the element that you are hovering on, no guesswork here.
Edits are rules, not mutations. Changing the styling of elements never touch the element; it compiles into a CSS rule that is injected into a stylesheet, precedence is worked out through an algorithm that considers the current cascade of that element’s styles. This way, even if there are cases where the iframe gets unmounted, I can keep re-applying the stylesheet rules and arrive at the same state idempotently without any other patch-work mechanism. The implementation of this module was really smooth when building out the whole feature. There is one edge case though, text, which had to mutate the DOM with an exceptional function with its own set of guards to ensure mutations are not lost.
Finally, since we stamped everything on build time, we had the exact location of the elements in the source code and passed all of it along to the Figr agent for cheap and exact edits.
8. conclusion
I will not lie; opening 1001 screens on a 4 GiB machine still dies when everything does not match up. A sudden zoom-out on a full board can glitch while the deferred layout of hundreds of nodes comes due at once; the real fix is a single composited layer at far zoom, which is underway.