Gleam doesn't compile to Erlang source anymore | Gleam programming language
Pangram verdict · v3.3
We believe that this entire text is human-written.
AI likelihood · overall
HumanArticle text · 1,600 words · 1 segments analyzed
Gleam is a type-safe and scalable language for the Erlang virtual machine and JavaScript runtimes. Today Gleam v1.19.0 has been published, so let's go over what's new. A new compilation target Over the last few months Giacomo Cavalieri has entirely rewritten Gleam's Erlang code generator that has an entirely different design, and most notably, outputs a different format. Previously Gleam generated Erlang source code, now it generates Erlang abstract forms. "Erlang abstract forms" is an intermediate representation used by the Erlang compiler. It is a metadata-annotated tree that represents Erlang syntax, and it is normally produced by running Erlang's tokeniser and parser. It has a binary encoding using Erlang's external term format, and with this binary format we can load our generated code directly, skipping the front-half of the Erlang compiler. This new Erlang code generator brings several benefits: The performance of the compiler has been improved, significantly reducing the build times for Gleam projects running on Erlang. The location metadata available to the runtime is now accurate to original Gleam source code, rather than to the Erlang source code the compiler would generate. This means, for example, the line numbers in BEAM crash reports and stacktraces are perfectly accurate, while previously they could be inaccurate, only pointing to the nearest function. This metadata could also enable full support for Gleam in debuggers such as edb, though we have not done any work on this ourselves. The code-quality of the Gleam compiler has been improved. The Erlang code generator was one of the oldest and most stable parts of the Gleam codebase, so while it wasn't causing us any problems it wasn't conforming to the standards and conventions we have today. This new replacement is excellent, and arguably raising the bar for the compiler as whole. We never have to hear someone use the word "transpiler" as a pejorative ever again.1 Ok, so how fast is it? I'm going to show you some numbers in a moment, but please remember that benchmarks are always contrived and never tell the full story. This data could be a good introduction or jumping-off point, but a good understanding requires the person to do further research and experience. The benchmark is based on José Valim's langcompilebench project. Thank you José! It is a measure of the time taken to compile 100 modules that each contain 100 functions that return a "hello world" string. This is practical as this shape of test project can be easily replicated across different languages to produce the most like-for-like test projects, but it is limited in what it can tell us as only a small subset of the features of each language get compiled. In real projects code will be greatly more varied, and different features will have different compilation costs in different languages. The first stage of the code generator rewrite was released in v1.18.0, the previous release, so let's compare v1.17.0 to the newly released v1.19.0. This chart shows the time taken to compile the benchmark project, lower is better. 0ms100ms200ms300ms400ms500msGleam v1.19Gleam v1.17 As you can see, a considerable improvement! This is a full build from scratch, without any caching. Gleam's compilation is incremental, so during typical development it would be much faster as it will not be compiling the entire project. The original langcompilebench includes only Erlang, Elixir, and Gleam, but I have extended it an assortment of other popular programming languages, to help folks get a rough feel for how fast Gleam compiles compared to a language they are familiar with. I've also included Gleam when compiling to JavaScript. Here's the results: 0ms200ms400ms600ms800ms1000msGleam (JavaScript)GoGleam (Erlang)ErlangJavaElixirElmRustC#TypeScript 7 Remember: This is a contrived benchmark and is this alone is insufficient to draw any hard conclusions about these languages. That said, these results do suggest that Gleam's compilation is nice and fast, and as a Gleam programmer I can say that Gleam development is very enjoyable, with little time spent waiting for the computer. Why not target BEAM bytecode directly? We have moved from from compiling from source code that is fed to the Erlang compiler to an intermediate-representation that is fed into the Erlang compiler, but why not bypass the Erlang compiler altogether? Couldn't we make a BEAM bytecode generator that outperforms the Erlang compiler? Perhaps we could also use Gleam's type information to generate more optimised code too. While it is possible to achieve these benefits, it's unlikely we would be able to. Unlike Erlang source and Erlang abstract forms, BEAM bytecode is not fixed and unchanging. Each new release of the virtual machine can evolve and improve the bytecode, adding new functionality and sometimes removing functionality that has been made redundant. We would need to commit to forever keeping up-to-date with this evolution, working closely with the Erlang maintainers to be ready for up-coming changes, and to have new versions of Gleam ready for new releases of the virtual machine. It would also be a significant effort to reproduce all the existing optimisations that have been implemented in the Erlang compiler over the decades, even with the help we might have from Gleam's more capable static analysis. Gleam is a community project supported by sponsorship. We have only a fraction of the finances of languages that are backed by corporations or academic institutions, so we need to think carefully about the most efficient and sustainable ways to use our resources. Gleam is a reliable foundation for software development, every decision we make has to work for years and decades to come. Compiling to Erlang abstract forms is the cost-benefit sweet-spot for Gleam today. We're also in great company with this decision. Our much-loved older-sibling language Elixir also compiles to Erlang via abstract forms. If it's good enough for Elixir, then it's good enough for Gleam! Alright, enough about that. There's plenty more in Gleam v1.19.0 to go-over. JavaScript decision tree assignment optimisation It's not just the Erlang code generation that has seen some love, there's some good improvements for JavaScript too. In Gleam flow control is done with pattern matching using a case expression, and it gets compiled to nested if statements. Because pattern matching is declarative the compiler is able to reorder and optimise the runtime logic, using a divide-and-conquer approach to find the right branch as quickly as possible. John Downey has improved this process to generate flatter code, with nested if statements collapsed into a single condition and fewer intermediate variables. For example, take this Gleam code: pub fn go(x) { case x { Wibble(1, 2) -> 1 _ -> 2 } } Previously this small bit of Gleam would compile to this surprisingly large bit of JavaScript2: export function go(x) { if (isWibble(x)) { let $ = x[0]; if ($ === 1) { let $1 = x[1]; if ($1 === 2) { return 1; } else { return 2; } } else { return 2; } } else { return 2; } } But now it generates this: export function go(x) { if (isWibble(x) && x[0] === 1 && x[1] === 2) { return 1; } else { return 2; } } A nice improvement, I'm sure you'll agree! Surprisingly this makes little-to-no change to the size of code bundles once minified and compressed (gzip really is magic), but the resulting code has fewer branches for JavaScript engines to optimise. Thank you John! List literal optimisation While they share a syntax in their respective languages, Gleam's immutable persistent list type is not the same as the JavaScript mutable contiguous array type. When Gleam code is compiled to JavaScript any list literal has to be compiled to JavaScript code that constructs the runtime data structures. For example, take this Gleam code: let numbers = [1, 2, 3] This would compile to JavaScript code2 like this, where a JavaScript array is constructed and passed to a function to convert it to a Gleam list. const numbers = arrayToList([1, 2, 3]) With this release the compile will now generate different code for short list literals, generating more direct code that does not convert from an array. const numbers = prepend(1, prepend(2, prepend(3, empty))) With modern JavaScript engines this results in a nice performance improvement, and it is especially impactful for projects that use lots of short lists, like those using the Lustre library. We recorded no improvement for longer lists, so the array-to-list approach is still used for those. Thank you Giacomo Cavalieri for this! TypeScript API overloads When compiling to JavaScript the Gleam compiler will also generate functions for working with the programmer-defined data structures from JavaScript. Alongside this the compiler can also provide TypeScript declaration files, enabling full integration between TypeScript and Gleam in a single project. One of the functions provided for each custom type is a function to check whether a value is a particular variant or not. For example, given the following type: pub type Box(value) { Full(value) Empty } The generated function would have this TypeScript declaration: export function Box$isFull(value: any): value is Full<unknown>; The keen-eyed TypeScript programmer readers may notice a problem here. If the value is already known to be of type Box<number>, then this function can be used to refine the container type to Full, but the type parameter of number is generalised to unknown, causing type information loss. This is very cumbersome. Giacomo Cavalieri has added an overload to the definition, so the type is preserved whenever possible. export function Box$isFull<I>(value: Box$<I>): value is Full<I>; export function Box$isFull(value: any): value is Full<unknown>; Thank you Giacomo!! Improvements for other build tools Gleam users will typically use the official build tool that is part of the