Superficie: Clojure you can read without learning Lisp

Syntactic sugar causes cancer of the semicolon.

Alan Perlis, Epigrams on Programming, 1982

Every Clojure example on this site opens in a tab called Superficie, with the original Clojure one click away. Perlis warned that syntactic sugar causes cancer of the semicolon, and Superficie is a lot of sugar, taken on one condition, that the two tabs show exactly the same program. This article explains why we built it, how that condition is kept, and how libraries teach it to write numerical kernels, proofs, and probabilistic models the way their fields do.

A few days is too long

Superficie began with a problem during a PhD in machine learning. The research code was written in Clojure, and the colleagues, supervisors, and domain experts who needed to read it worked in Python. Showing a function in a meeting, a paper, or a code review meant first explaining the parentheses.

That unfamiliarity passes. Most programmers read S-expressions comfortably after a few days. A meeting, a blog post, or a review with someone outside the team does not last a few days, though. The reader skips the code and trusts the prose around it, which defeats the purpose of showing code at all.

Our stack makes the problem larger. Datahike, Spindel, Yggdrasil, Raster, and Ansatz are Clojure libraries, and most people who read about them do not write Clojure. Simmis is meant to be usable without knowing Clojure, while deep work on the core stack still requires it. Superficie is the tool we use at that boundary. We keep writing Clojure, and we show it in a notation that readers of Python, Julia, TypeScript, or Lean can follow on first sight.

One program in two notations

Superficie is a bidirectional renderer. It prints Clojure forms as text with ordinary calls, infix arithmetic, and blocks, and it reads that text back into the same forms. Here is a function from our Stuttgart retail simulation that compares two shop names by the words they share:

What is this syntax?
defn jaccard [a b]:
  a := set(str/split(a, #" "))
  b := set(str/split(b, #" "))
  u := count(set/union(a, b))
  if zero?(u):
    0.0
  else:
    double(count(set/intersection(a, b))) / u
  end
end
(defn jaccard [a b]
  (let [a (set (str/split a #" ")) b (set (str/split b #" "))
        u (count (set/union a b))]
    (if (zero? u) 0.0 (/ (double (count (set/intersection a b))) u))))

A few rules cover most of what you see. A call is f(a, b). Arithmetic and comparisons are infix and need spaces around the operator, so a-b stays a Clojure name and a - b is subtraction. A block starts with a head and a colon and ends with end, which covers defn, if, let, loop, and every other form with a body. A let that ends a body is written as name := value statements, so a function reads from top to bottom instead of nesting one level deeper per binding.

To try it, paste any Clojure into the playground, which renders it live and includes a SCI REPL that evaluates Superficie in the browser.

Superficie has no runtime of its own. Any Clojure source prints as Superficie. Superficie source can also be read and evaluated directly, on the JVM, in Babashka, or in the browser through a SCI REPL, and it works with every Clojure library because what runs is ordinary Clojure. Macros and syntax-quote print as themselves rather than as their expansion, so a macro definition reads like the template its author wrote.

The same program, or it is a bug

A rendering that is almost right is worse than none, because a reader cannot tell which part changed. Superficie’s rule is exact roundtripping. We print the forms, read the text back, and compare the result with the original forms for equality. The comparison treats as equal only spellings that Clojure evaluates identically, such as the names Clojure’s reader invents for the parameters of #() or (new Foo x) against (Foo. x).

We check the rule against real code rather than hand-picked samples. In the current version, 421 of the 434 readable files in Raster and Ansatz roundtrip exactly, both in the compact form and in the width-aware layout used on this site. Across 13 open source Clojure projects, among them Datahike, DataScript, Malli, SCI, Babashka, and core.async, 622 of the 670 files our checker could read do. All 41 files of the Stuttgart simulation do.

Most of the remaining files use one of Superficie’s block words, such as match, let, ns, or end, as an ordinary name in a position where Superficie reads a block. The printer avoids this where it can. A local named end keeps its let block instead of becoming end := now(), which would close the enclosing block, and a statement that starts with such a word is parenthesised. Where the printer cannot avoid it, the constraint is documented, and code written in Superficie from the start does not run into it.

The corpora keep finding cases that a design review misses. A proxy block followed by a vector on the next line read its closing end as the name of a method. In the JavaScript build that renders this site, 4.0 printed as 4, because JavaScript has a single number type, and a reader of a numerical kernel would have seen integer arithmetic. Both are fixed, and the checker found both in code nobody wrote to test it.

The printer never fails, either. A form it has no block for prints as a call, f(a, b, c), which always reads back. That fallback is exact and hard to read, so most of the recent work went into making it rarer.

Where the notation comes from

Lisp was meant to have a second notation from the start. John McCarthy’s 1960 paper on Lisp wrote programs as M-expressions, such as f[x; y], and used S-expressions to represent data. In his History of Lisp, he recalled that translating M-expressions into S-expressions “was neither finalized nor explicitly abandoned. It just receded into the indefinite future, and a new generation of programmers appeared who preferred internal notation to any FORTRAN-like or ALGOL-like notation that could be devised.” Later attempts include Dylan, which replaced its prefix syntax with an Algol-like one in the 1990s, and the readable project’s SRFI 105 and SRFI 110 for Scheme. The most thorough recent one is Rhombus, a language built on Racket and described by Matthew Flatt and colleagues at OOPSLA 2023.

For a Clojure programmer, the interesting part is what a surface syntax has to rebuild. In an S-expression the parentheses are the tree, so Clojure’s reader stays small and knows nothing about operators. A notation with infix operators and blocks has to recover that tree in two steps, and Superficie takes both from existing work. First, following Rhombus’s shrubbery notation, it groups the text by its brackets without knowing what any operator means. A mismatched bracket becomes an error node in that tree while the rest of the file still parses, which is what makes useful error messages possible. Second, it decides precedence with Vaughan Pratt’s 1973 technique: each operator has a binding power, and the parser keeps consuming operators while their power exceeds the current minimum, so a + b * c groups correctly without a grammar rule per precedence level.

The visible choices are borrowed where readers already know them. A block header ends with a colon, as in Python, and a block closes with end, as in Julia, so indentation carries no meaning and copying code through HTML, chat, or a diff cannot change it. Arrow lambdas and indexing follow Julia, pattern-match arms and := follow Lean, and stores follow OCaml. Operators need spaces around them because Clojure names may contain -, *, ?, and <.

Libraries decide how their macros read

The Rhombus paper observes that “the message of macros has been difficult to detangle from Lisp’s minimalistic, parenthesis-oriented notation.” Languages without parentheses answer it in two ways, and Superficie takes the second.

In Rhombus, parsing is not finished before macros run. The reader leaves a sequence such as 1 + 2 * 3 flat, and the expander completes it while expanding macros, so a library can define a new operator together with its precedence. In Julia, the grammar is fixed. The parser builds the whole tree first, a macro call is marked with @, and the macro receives that tree as data and returns a new one.

Superficie parses first, as Julia does, but the data a macro receives is plain Clojure. a + b arrives as (+ a b) and a block arrives as the list it stands for, so a macro written for Clojure works unchanged and never sees Superficie. No @ is needed either. A macro call is written like any other call, and Clojure decides at expansion time that it names a macro. Superficie gives up Rhombus’s user-defined operators and gets every existing Clojure library in exchange. A macro can be written in Superficie too, with syntax-quote as its template:

What is this syntax?
defmacro unless [test & body]:
  `if not(~test):
     do(~@body)
   end
end

unless(ready?(job), log("waiting"), retry(job))
(defmacro unless [test & body]
  `(if (not ~test) (do ~@body)))

(unless (ready? job) (log "waiting") (retry job))

Most of our interesting code lives inside library macros. Raster’s deftm defines a typed function that Raster can compile into a GPU kernel, Ansatz’s a/theorem states and proves a theorem, and Spindel’s spin wraps a program such as the probabilistic model of the Stuttgart simulation. Without more information, Superficie can print a macro only as a call, deftm(weight, [att :- Double], :-, Double, ...), which is exact and unreadable.

A shape tells Superficie how a macro’s arguments split into a block header and a body. For unless, the shape [:form :body] lets unless(ready?(job), log("waiting"), retry(job)) read as unless ready?(job): followed by its body and end, and the macro still receives the same list. The shape changes the notation, never what the macro sees. Shapes for Raster, Ansatz, and Spindel ship with Superficie. A library can declare its own, in a resource file, in the macro’s metadata, or at runtime. The printer uses a shape only after checking that the block reads back to the original form, so a wrong shape can cost readability but not correctness.

A shape can also change how a block’s body is written. This kernel from the Stuttgart simulation computes, for every 100 m cell, the normaliser of the shop choice distribution:

What is this syntax?
deftm cell-totals!
    "Z_c for every cell: the normaliser of the choice distribution."
    [cell-lon :- Array(double),
     cell-lat :- Array(double),
     cand-lon :- Array(double),
     cand-lat :- Array(double),
     cand-att :- Array(double),
     params :- Array(double),
     totals :- Array(double),
     n-cells :- Long,
     nc :- Long] :- Void:
  alpha := params[0]
  beta := params[1]
  d0 := params[2]
  par/map-void! c n-cells:
    lon := cell-lon[c]
    lat := cell-lat[c]
    loop [q int(0), z 0.0]:
      if q < nc:
        recur(unchecked-add-int(q, 1),
          z + weight(cand-att[q],
                haversine-m(lon, lat, cand-lon[q], cand-lat[q]), alpha, beta, d0))
      else:
        totals[c] <- z
      end
    end
  end
end
(deftm cell-totals!
  "Z_c for every cell: the normaliser of the choice distribution."
  [cell-lon :- (Array double), cell-lat :- (Array double),
   cand-lon :- (Array double), cand-lat :- (Array double), cand-att :- (Array double),
   params :- (Array double), totals :- (Array double),
   n-cells :- Long, nc :- Long] :- Void
  (let [alpha (aget params 0) beta (aget params 1) d0 (aget params 2)]
    (par/map-void! c n-cells
      (let [lon (aget cell-lon c) lat (aget cell-lat c)]
        (loop [q (int 0) z 0.0]
          (if (< q nc)
            (recur (unchecked-add-int q 1)
                   (+ z (weight (aget cand-att q) (haversine-m lon lat (aget cand-lon q) (aget cand-lat q)) alpha beta d0)))
            (aset totals c z)))))))

Raster declares that inside its kernels x[i] means aget and x[i] <- v means aset. That declaration names functions; it does not fix what indexing means. Raster’s aget and aset dispatch on the array’s element type, much as Julia’s getindex and setindex! do, so the brackets are as polymorphic as the functions behind them. Outside a block that declares indexing, aget(U, i) stays a call, and in a block header name[x] is never read as an index.

Ansatz uses the same mechanism for a different field. Its terms follow Lean, so inside a/defn and a/theorem a name like Nat.succ(n) is a call to a dotted name rather than a Java method call, and a pattern match prints as Lean-style arms:

What is this syntax?
a/defn rb-size [t :- RBTree(Nat)] Nat:
  match t:
    | leaf => 0
    | node(color, left, key, right) => 1 + (rb-size(left) + rb-size(right))
  end
end

a/theorem map-preserves-len [f :- arrow(Nat, Nat) l :- List(Nat)] (llen(lmap(f, l)) = llen(l)):
  induction(l)
  all_goals(grind("lmap", "llen"))
end
(a/defn rb-size [t :- (RBTree Nat)] Nat
  (match t
    [leaf 0]
    [(node color left key right) (+ 1 (+ (rb-size left) (rb-size right)))]))

(a/theorem map-preserves-len [f :- (arrow Nat Nat), l :- (List Nat)]
  (= (llen (lmap f l)) (llen l))
  (induction l) (all_goals (grind "lmap" "llen")))

Choosing the notation

Each shorthand had to meet two conditions. A reader who has never seen Superficie should recognise it, and it must read back as exactly one form. Several candidates met the first condition and failed the second, and those failures shaped the syntax.

Anonymous functions have two spellings. A one-line function passed as an argument prints as an arrow, map(x -> x * x, xs), and its body runs to the next comma or closing bracket. The arrow needs a space on each side. Without them, a->b is one Clojure name, and base ->(raw, f()) is a name followed by a call to Clojure’s thread-first macro, which is how it appears in a let binding. Clojure’s #() literal keeps its own syntax with % parameters. By the time Superficie sees Clojure source, the reader has already expanded #(inc %) into a function whose parameter is called something like p1__123#. Only the #() reader produces names of that pattern, so Superficie can recognise them exactly and print #(inc(%)) again instead of the expansion.

What is this syntax?
defn summarize [xs]:
  {:squares map(x -> x * x, xs),
   :total reduce((acc, x) -> acc + x, 0, xs),
   :labels map(#(str("item-", %)), xs)}
end
(defn summarize [xs]
  {:squares (map (fn [x] (* x x)) xs)
   :total (reduce (fn [acc x] (+ acc x)) 0 xs)
   :labels (map #(str "item-" %) xs)})

Indexing was the harder decision. Brackets for Clojure’s general get would read naturally in data code, but x[i] would then have two meanings depending on where it appears, and get returns nil for a missing index where aget fails. Keyword lookup such as (:name user) is also several times more common than get in a rough count over the projects above, so a global index syntax would mostly dress up a less common idiom. Indexing therefore stays with the libraries that declare it.

Stores needed an operator that nothing else could be mistaken for. = is equality in Superficie, as in Clojure, and := introduces a binding. A store is neither, so it is written <-, as OCaml writes a.(i) <- v. Binding, comparison, and mutation each have one spelling.

The := statements are the newest rule. A let in the last position of a body prints as statements, consecutive statements read back as one let, and a let used as a value keeps its block. The printer declines to flatten when flattening would change the form. When a let’s only body is another let, the two would read back as one, so the inner one stays a block.

Where we use it

In Simmis, a Code View setting shows code in Superficie instead of Clojure, read-only. It applies in chat, in documents, and in the run inspector, where a person reviews the code an agent ran. Simmis is meant to be usable without knowing Clojure, and this is how someone who does not read Clojure can still check what an agent did.

On this site, Superficie renders every Clojure example at build time through its npm package, and the Superficie tab is shown first. An article can name the requires its snippets assume, so an excerpt without its namespace form still shows Raster kernels and Ansatz theorems as blocks. datahike.io uses the same package for its documentation. The project is on GitHub under the Apache 2.0 licence.

From reading code to changing it

Everything above concerns reading. Two open questions concern changing code, and we have not answered either.

The first is editing. Superficie reads back exactly the forms it printed, but not the comments and line breaks the author chose, because Clojure’s reader discards comments before Superficie sees the code. Editing a Clojure file in Superficie and writing Clojure back would need both preserved, so that a Clojure programmer still recognises the result as their own file. Until then, Superficie is a way to show existing code and to write new code, not to edit existing Clojure files.

The second is review. Simmis already shows the code an agent ran in Superficie. A proposed change is harder to show, because a line diff of the Clojure text does not map line by line onto Superficie. A diff over forms could, and it raises its own questions, such as how to show a changed argument inside an unchanged block and how to recognise a definition that moved. Both questions matter for the reason this project exists. More of the people who decide whether a code change should be adopted will not read Clojure.

Whatever the answers, the rule the rest of this article depends on stays. Superficie reserves its block words, a snippet without its namespace needs the requires it assumes, and a rendering that does not read back as the same forms counts as a bug.

FROM AGENT OUTPUT TO ORGANIZATIONAL WORK

See how Simmis lets teams delegate consequential work without losing control of what becomes official.

build with simmis Simmis Cloud waitlist