hecc - hachem's experimental c compiler

hecc (pronounced: heck), is an experimental C compiler built around a relatively simple idea: C often does not give the compiler enough information about what the programmer intended. A pointer assignment might be an alias, an ownership transfer, or simply an arbitrary operation whose meaning is known only to the programmer. C gives the compiler the resulting values, but usually not the intent behind them.

hecc explores whether a small number of explicit extensions can let the programmer write that intent down, so the compiler can check it and build on it. The point is not to turn C into Rust, C++, or another large systems language, and it is a language extension you write, not an analyzer you point at code you leave unchanged: existing C keeps compiling untouched, and where you reach for the extensions -- ownership transfers, borrows, tagged unions -- you gain expressiveness and checking the base language cannot offer. The question is how much a small, additive layer can add without giving up what makes C worth using.

The nearest one-line description is gradual typing, aimed at memory responsibility rather than at types. Like TypeScript over JavaScript, the checks are opt-in, adopted a file at a time, unsound on purpose, and escape-hatched -- useful precisely because they describe intent the programmer states rather than intent a tool has to guess. Nothing here is a feature C-safety work has not tried before; the bet is on that combination held to a hard rule -- everything erases to ordinary C, nothing is ever required, and the code is never trapped inside the toolchain. To keep the "small" claim honest, the document is in two parts: a deliberately small core that is what hecc actually claims to be, and, past a divider, a set of explorations held to the same discipline that have not yet earned their way in.

The project is experimental, and much of the semantics described here are still being worked out. This is not intended to be a finished language specification. It is a document for keeping track of the direction of the project and the reasoning behind its features while those features are still being explored.

Why C?

C has a number of problems, but its size is not one of them. The language itself is relatively straightforward to learn, its standard library is comparatively small, and people regularly replace parts of its ecosystem with their own implementations. Hobby operating system developers write their own allocators, containers, runtimes, standard libraries, and sometimes entire compilers. There are countless alternative C libraries because the language is small enough that understanding and replacing its pieces is realistic.

That experience changes with larger systems languages. C++, Rust, and Zig all solve problems that C does not, but they also introduce significantly larger languages and ecosystems. The goal of hecc is not to gradually reproduce those languages inside C. If every problem is solved by adding another type, annotation, keyword, library component, or compiler feature, the language will eventually become exactly the kind of thing it was intended to avoid.

The question behind hecc is therefore not how to make C completely safe. It is how much additional information can be given to the compiler while keeping the language recognizably C.

How hecc Is Built

hecc is, at least at first, a transpiler: it parses the extended syntax, runs its analysis, and emits ordinary C for an existing compiler to build. This matters mainly for how the rest of this document reads. Most of hecc's extensions are information for the analysis stage and nothing more -- once checked, they erase, and the emitted C looks like what a C programmer would have written by hand. Where a later section says an extension "generates" some C or "erases at code generation", this is the step it means. The full pipeline, the LLVM backend, and what the transpile-to-C design buys a project are covered under Transpilation.

Additive by Design

Everything in hecc is meant to be strictly additive. Using a feature gains you something; not using it costs you nothing, and existing C compiles and runs unchanged. There is no point at which adopting hecc first requires giving something up.

Most tools that improve on C are subtractive before they are additive. Rust asks for a rewrite, static analyzers ask you to wade through false positives, sanitizers ask for runtime. Each takes something before it gives anything back. hecc's floor is ordinary C at no penalty, and every feature sits above that floor as an opt-in gain.

The ownership system is the clearest example. = is exactly C's assignment -- it makes no ownership claim and is never required to be checked, so the aliasing and borrowing that existing code is full of keeps working untouched. <- and the annotations built on it are the opt-in: reach for them where the guarantee is worth the keystrokes, and leave = everywhere you do not care. Memory safety becomes something you add, one line at a time, rather than a gate to clear before the code will build. The same holds for the rest -- type unions, bounds annotations, the rulesets -- none are required, and none change what unannotated code means.

This is a firm rule. A feature that made existing C warn, fail, or need editing to adopt would break it, and does not belong in hecc. hecc can reward you for saying more; it never penalises you for saying nothing.

Opt-in is sometimes read as a weakness next to a language that is safe by default, but that is a difference in compilation model, not in reachable safety -- and opting into strictness is already how C is written. Almost nobody ships serious C without -Wall -Wextra -Werror, a sanitizer, or a MISRA-style ruleset; reaching for extra checking is the norm, not the exception. hecc's features are the same reflex. The intended end state of a hecc codebase is to have adopted the safety layer, not to sit on the bare floor forever: writing good hecc means using <-, origins, and the strict and runtime tiers. Which of those a build enforces is a policy the project sets -- strict for the code it wants fully checked, compatibility for the C it is still adopting -- the same way it sets -Werror, so "good hecc" is enforced rather than merely hoped for. How those defaults are chosen is covered under Strictness and Compatibility. This is no different from any language having a notion of correct use. You can write Rust that leaks -- allocate and mem::forget, or build a reference cycle -- and it compiles, but it is not correct Rust. = everywhere is legal hecc in the same spirit: permitted, adoptable, and not the goal.

Ownership

The main feature currently being explored in hecc is an ownership system for pointers. Consider ordinary C:

int* a = malloc(...);
int* b = NULL;

b = a;

At runtime, this is straightforward. Both a and b now contain the same address. The problem is that the compiler does not know what the assignment means. Did the programmer create another reference to the allocation? Is b now supposed to be responsible for freeing it? Is a still responsible? Was this assignment accidental? C cannot answer these questions because the assignment itself does not communicate that information. hecc experiments with making ownership transfers explicit:

int* a = malloc(...);
int* b = NULL;

b <- a;

The <- operator means that ownership of the allocation is being moved from a to b. This does not need to change the runtime behaviour of the program. Once hecc has checked that the operation is valid, <- erases and the emitted C is simply:

int* a = malloc(...);
int* b = NULL;

b = a;

The difference exists entirely during compilation. The generated program is still fundamentally ordinary C. The important distinction is that <- communicates something that = does not. = copies a pointer without saying anything about ownership. <- explicitly tells the compiler that responsibility for the allocation is moving.

Owners

When memory is allocated through something the compiler understands, such as malloc(), hecc can establish an owner:

int* a = malloc(...);

At this point, a owns the allocation. That ownership can eventually be released:

free(a);

or transferred:

int* b = NULL;

b <- a;

After the transfer, b owns the allocation.

The exact pointer values do not necessarily change. a and b may both still contain the same address. Ownership is not a property of the address itself; it is information tracked by the compiler about which pointer is currently responsible for the allocation.

Conceptually, before:

a -> allocation [owner]

after:

a -> allocation
b -> allocation [owner]

The original pointer does not necessarily become invalid as a pointer. It simply no longer owns the allocation. This is an important distinction. hecc is not attempting to prevent pointer aliases. C programs constantly use multiple pointers referring to the same memory, and preventing that would make the language substantially less useful for the kind of programming it is intended for.

Receiving Ownership

A pointer should not be able to silently overwrite existing ownership:

int* a = malloc(...);
int* b = malloc(...);

b <- a;

If both pointers own separate allocations, transferring ownership into b would leave its previous allocation without an owner. The simplest rule is that a pointer receiving ownership must not already own an allocation. A typical transfer would therefore look like:

int* a = malloc(...);
int* b = NULL;

b <- a;

b is explicitly initialized and does not currently own anything, so it can receive ownership from a. This is one of the places where the system deliberately asks the programmer to be explicit. Ownership should not be accidentally overwritten. There may eventually be other valid ways for a pointer to receive ownership, but the basic rule provides a clear starting point: ownership moves into a pointer that does not currently have responsibility for another allocation.

Ordinary Pointer Copies

hecc should not make ordinary C pointer assignments illegal. This remains valid:

int* a = malloc(...);
int* b = a;

The programmer has copied the address and created an alias. That is ordinary C and should continue working. However, this is not necessarily the correct way to express a tracked relationship between two references. The compiler now has two pointers referring to the same allocation:

a ----\
       -> allocation
b ----/

The compiler may know that a originally owns the allocation, but the relationship between the pointers has not been explicitly described. If a transfers ownership somewhere else, b still exists. If a is freed, b may become dangling. If either pointer is returned from a function, stored somewhere else, or passed through several layers of code, the compiler can lose track of which pointers belong together. Ordinary pointer copies should therefore remain allowed, and a copy like this is treated as a borrow of whatever a owns. Keeping that relationship intact once the reference escapes the function is what borrows and origins are for, described shortly.

Ownership Through Functions

Functions introduce another important ownership case. If hecc can inspect a function's implementation, it can analyse how pointers are used and infer ownership behaviour where possible. For example:

int* create_buffer(void)
{
    int* buffer = malloc(...);

    return buffer;
}

The compiler can see that buffer receives ownership from malloc() and that the allocation is returned from the function. The caller therefore receives ownership:

int* buffer = create_buffer();

External functions are different. If hecc cannot inspect the implementation, it cannot automatically know what happens to a pointer passed to the function. The function might borrow it, modify the memory it points to, store the pointer somewhere, free it, or take ownership of the allocation. For cases where ownership is transferred into a function, hecc can use the same <- syntax used for normal ownership transfers:

function(<- pointer);

This explicitly tells the compiler that ownership of pointer is being moved into function. After the call, the function is responsible for the allocation rather than the caller. Ordinary function calls continue to mean that ownership is not transferred:

function(pointer);

The pointer is passed normally, and the caller retains ownership unless the compiler can determine otherwise from the function's implementation. This keeps the ownership syntax consistent. <- always represents an explicit ownership transfer, regardless of whether the destination is another pointer:

other <- pointer;

a function:

function(<- pointer);

or the caller, when a function hands ownership back out through its return:

return <- pointer;

return <- pointer; is the counterpart to function(<- pointer);. One moves ownership explicitly into a function, the other moves it explicitly back out to the caller. When hecc can inspect a function it can usually infer this on its own, exactly as it did for create_buffer() above, so the explicit form is mainly for functions the compiler cannot see into or for cases where the programmer wants the transfer to be unambiguous.

The preferred approach is still to avoid annotations where inference is possible. If hecc can inspect a function and determine what happens to a pointer, additional syntax should not be necessary. Explicit ownership syntax is primarily useful when the compiler cannot know what a function does or when the programmer wants to make the transfer unambiguous.

Consuming Ownership

Passing ownership into a function at the call site with function(<- pointer) works, but it puts the contract in the wrong place. Every caller has to remember to write it, and the function's own signature says nothing. A function whose whole job is to take ownership -- a destructor, a close that owns its handle, a container that adopts an element -- should declare that once, on itself. hecc lets the parameter carry the transfer:

void destroy(struct Thing* <- t);

The <- on the parameter says destroy takes ownership of t. After the call the caller is no longer the owner; destroy is responsible for freeing it or handing it on. The contract now lives on the signature, so every caller is checked against it and the mistakes that follow from a consumed pointer can be flagged:

struct Thing* t = create();

destroy(t);   /* ownership consumed here */

use(t);       /* warning: t was consumed by destroy */
free(t);      /* warning: t was consumed by destroy -- double free */

Because the signature already declares the transfer, the call does not repeat <-; the compiler knows every call to destroy consumes its argument. Writing destroy(<- t) is still allowed where the hand-off should be visible at the call site, and for an external function whose signature is not annotated, the call-site <- remains the only way to say it.

Round-Trip Ownership

A function can also take ownership and give it back. If it hands ownership back through its return value, no new operator is needed -- that is just a consuming parameter and a return <-:

struct Node* rebalance(struct Node* <- tree)
{
    /* ... */
    return <- tree;   /* or a different node that now owns the tree */
}
struct Node* tree = build();

tree <- rebalance(tree);   /* consumed by the call, re-owned from the return */

tree is consumed by the call, so it owns nothing at the assignment and can receive ownership again from the return without breaking the rule against overwriting ownership it still holds. A dedicated operator for this would be redundant.

<-> is for the case those pieces cannot express: when ownership comes back through the parameter rather than the return, leaving the return value free to carry something else. That means passing the owning pointer by reference -- a pointer to it -- which the function writes back:

bool grow(int** <-> buffer, size_t extra);

<-> says ownership of *buffer passes into grow and comes back out through *buffer. The function may free or reallocate the allocation and writes the owning pointer back through buffer; on return the caller owns again, possibly a different allocation. Because ownership returns through the parameter, the return value means whatever the function wants -- here, whether the growth succeeded:

int* buffer = malloc(...);

if (!grow(&buffer, 128))
    ; /* handle failure; buffer still owns whatever grow left it holding */

use(buffer);   /* buffer owns again -- no re-assignment from a return needed */

This is the in/out owner parameter, and it is the one shape the other markers cannot state: <- alone consumes and never gives ownership back, return <- sends it out through the return, and <-> is what says it left through the argument and came back through it. realloc is a related case where the transfer is conditional -- on failure the original stays owned -- which is why hecc models it specifically; that is covered under realloc() below.

Conditionals

Ownership becomes significantly more complicated once control flow is involved. Consider:

int* a = malloc(...);
int* b = NULL;

if (condition)
    b <- a;

Before the conditional, a owns the allocation. Inside the conditional, ownership is transferred to b. After the conditional, the compiler has two possible ownership states:

condition == true:  b owns the allocation
condition == false: a owns the allocation

Neither a nor b can simply be described as the owner after the if without considering both possible execution paths. This becomes increasingly complicated with more control flow:

if (condition_a)
    b <- a;
else if (condition_b)
    c <- a;

Now the owner may be a, b, or c, depending on how execution proceeded. A flow-sensitive pass can follow these states through the control-flow graph and require the paths to reconcile at the join, which handles the case while the references stay in view. Once ownership escapes -- through a pointer that is stored in a structure, returned, or aliased across a translation-unit boundary -- a plain flow-sensitive pass loses the thread. Reconnecting the reference to what it depends on is what borrows and origins, described next, are for.

Borrows and Origins

The one problem ownership leaves open is what happens to a reference once it escapes the scope that created it. A pointer can be returned from a function, stored in a structure, or passed through several layers of code, and at that point the flow-sensitive analysis can no longer see the relationship between the reference and the allocation it depends on. An earlier version of hecc tried to solve this with symmetric "tracking groups" that marked a set of pointers as related. That was the wrong shape. The relationship is not symmetric, and "related" is not a checkable invariant. The relationship is directional -- one pointer borrows from another that owns -- and the borrowed pointer must not be used after the owner's responsibility ends. It is also the part of hecc whose vocabulary is most borrowed from Rust, which -- as the end of this section spells out -- makes it the easiest to mistake for Rust and the most important to keep distinct.

Borrows

hecc already distinguishes two kinds of pointer assignment: <- moves ownership, and = copies a pointer without transferring it. The pointer produced by = is a borrow: a reference that does not own the allocation and depends on some owner that does.

int* a = malloc(...);   /* a owns */
int* b = a;             /* b borrows from a */

Within a single function this needs no annotation. hecc can see that a owns the allocation and that b was copied from it, so it treats b as a borrow of a and can already warn if b is used after a is freed or moved:

int* a = malloc(...);
int* b = a;

free(a);
*b = 5;   /* warning: b borrows from a, which has been freed */

The borrow is still an ordinary int* at runtime. The borrow relationship is information for the analysis and erases like ownership itself; the generated C is unchanged.

Origins

A borrow becomes hard to track only when it crosses a boundary the analysis cannot see through -- a function signature or a structure definition. To carry the relationship across such a boundary, a borrow is tied to an origin: a name for the owner, or the owning scope, that the borrow depends on. The term follows the "origins" framing of Rust's Polonius borrow checker; it fits hecc better than the older word "lifetime" and avoids the confusion that led hecc to drop that word earlier.

An origin is written with a $ sigil. It names a relationship, not a variable, and it appears only where inference cannot reach: on signatures and on struct fields. The invariant it carries is the one the old tracking groups lacked:

A borrow must not be used after its origin ends. An origin ends when the owner it names is freed, moved elsewhere with <-, or goes out of scope.

The freeing case appeared above. A move ends an origin the same way, because ownership has left the pointer the borrow was tied to:

int* a = malloc(...);
int* b = a;         /* b borrows from a; a is its origin */

*b = 1;             /* fine: a still owns the allocation */

int* c = NULL;
c <- a;             /* ownership moves out of a -- a's origin ends */

*b = 2;             /* warning: b used after its origin ended */
*c = 3;             /* fine: c owns the allocation now */

Going out of scope is the third case: if the owner is a local that leaves scope, a borrow that outlives it is using an origin that has ended. Everything else is the machinery for propagating that invariant across boundaries.

Borrows across functions

When a function returns a pointer that borrows from one of its arguments, the return value's origin is that argument:

$a char* skip_spaces($a char* s);

The origin $a ties the returned borrow to s. A caller that frees the underlying buffer before using the result is using the borrow after its origin ended, and hecc can say so:

char* line = malloc(...);
char* rest = skip_spaces(line);

free(line);
puts(rest);   /* warning: rest borrows from line, which has been freed */

Used in the other order, while the origin is still alive, the same call is fine:

char* line = malloc(...);
puts(skip_spaces(line));   /* fine: line still owns here */
free(line);

As in Rust, most signatures do not need this spelled out. When a function has a single pointer parameter and returns a borrow, the output borrows from that parameter and the origin is inferred. The annotation is only required when there is more than one candidate and the compiler cannot choose:

$a char* longer($a char* x, $a char* y);

Here the result borrows from both x and y, so it is valid only while both are -- its origin is the shorter of the two.

The annotation earns its place at the call site, because it lets a caller be checked without the compiler ever seeing the body of longer:

char* a = malloc(...);
char* b = malloc(...);

char* longest = longer(a, b);   /* borrows from both a and b */

free(a);
puts(longest);   /* warning: longest borrows from a, whose origin has ended */

The signature alone carries enough to catch this. Freeing b instead of a would be flagged the same way, and freeing neither until after the last use of longest is fine. Without the $a on the signature, longer would look like any other function returning a char*: hecc would have no reason to connect the result to the arguments, the call would stay best-effort, and the misuse would pass unnoticed. That contrast is the whole point of the annotation -- it is the difference between a checkable boundary and an opaque one.

Function calls that pass a pointer with no origin, and external functions the compiler cannot inspect, stay best-effort exactly as before: if the relationship is neither stated nor inferable, hecc does not invent one.

Borrows in structures

A structure that stores a pointer must say whether it owns that pointer or borrows it, because the two have opposite consequences when the structure is destroyed.

A field the structure owns is freed when the structure's ownership ends, which lets ownership nest:

struct Buffer
{
    int* data;   /* owned: released when the Buffer is */
    size_t len;
};

A field the structure borrows ties the structure to an origin, and an instance must not outlive it:

struct Slice
{
    $a char* data;   /* borrows from origin a */
    size_t len;
};

A struct Slice now carries the origin $a. Building one binds $a to the borrowed buffer, and hecc checks that the slice does not outlive it:

char* buffer = malloc(...);
struct Slice s =
{ 
    .data = buffer, 
    .len = n
};

free(buffer);
use(s.data);   /* warning: s borrows from buffer, which has been freed */

Owned fields extend ownership inward; borrowed fields extend origins outward. Nested, self-referential, and mixed-ownership structures still need careful rules, but these two field kinds are the foundation the earlier discussion of structures was missing.

Conditionals again

Origins also settle the conditional case from earlier. When a borrow's validity differs between paths, it is valid after the join only if it is valid on every path, and its origin after the join is the shortest of the possibilities. This is the reconciliation the ownership analysis already performs, now with a defined relationship to reconcile rather than a guess.

Not a Borrow Checker

The vocabulary here -- ownership, borrows, moving, origins -- is taken from Rust, which makes it easy to assume hecc is doing what Rust does. It is not. The two share a handful of words and almost none of their substance, and they are better thought of as different tools that happen to reuse the same nouns.

Rust's borrow checker is a proof system. Its purpose is to make unsafe programs unrepresentable: code that cannot be shown safe does not compile. To do that, one mutable borrow is exclusive of all others, a move invalidates its source, and ownership is the spine the whole language rests on -- Drop, Send, Sync, and the type system are all built on it. You shape your program around what the checker will accept.

hecc's system is a description. Its purpose is to let the programmer state intent the compiler would otherwise have to guess, so it can flag likely mistakes. It proves nothing and forbids almost nothing. The details invert Rust's point for point:

  • Purpose. Rust enforces; hecc describes. One is a gate, the other is a note in the margin.
  • Default. Rust rejects anything it cannot prove safe. hecc accepts everything and warns only on what it can prove wrong. Opposite polarity.
  • Moves. A Rust move kills the source -- using it again is a compile error. hecc's <- transfers only the responsibility to free; the source stays a usable pointer, because C depends on aliasing.
  • Aliasing. Rust's exclusivity rule, one &mut exclusive of all &, is load-bearing: it buys soundness, data-race freedom, and noalias. hecc has no such rule and chases none of those. It allows any number of live borrows, mutable included, and tracks exactly one property: whether a borrow outlives its origin.
  • Place in the language. In Rust, ownership is the substrate and cannot be opted out of. In hecc it is an overlay on ordinary C that erases completely -- like ownership and bounds, and unlike type unions -- and can be ignored entirely. The language underneath is C, not a new model.

The cost of not having Rust's exclusivity is that hecc cannot promise the absence of data races or use borrows to justify aliasing optimizations. The gain is that borrowing can be added to existing C without restructuring how the code shares memory. So the parallels are real but shallow, and reading hecc as "Rust with the checking turned down" gets it backwards. Rust decides whether a program is allowed to exist; hecc reads a program you already wrote and tells you what looks off. They should be conceptualized as different things that borrowed the same words.

Strictness and Compatibility

How much hecc enforces is a setting, not a fixed property of the language, and the useful unit for it is the file or region rather than the whole program. There are two ends:

  • Compatibility -- plain C. = and its escapes are best-effort: analysed and warned about only where a problem can be proven, never required to be resolved. Existing C compiles and runs unchanged. This is what makes adoption free.
  • Strict -- every escaping = borrow must be resolved. A copy whose borrow leaves the scope has to be made a move (<-), given an origin, or explicitly marked untracked; a silent unresolved escape is an error, not a warning. Ownership, origins, and the leak checks are enforced rather than advisory.

Strict does not mean "no =". Plenty of = are legitimate borrows, and those stay legal -- strict only forbids an escape that is silent, one the compiler cannot account for and the programmer has not resolved. It closes the single real safety hole and nothing else.

The sensible default follows the code rather than a single global switch:

  • New code defaults to strict. Writing fresh hecc, you get the safety without ceremony, and "good hecc" is what you produce unless you opt out -- the same reason a new TypeScript project turns strict on.
  • Existing code defaults to compatibility. Dropping hecc onto a C tree builds, and you migrate file by file, flipping each to strict once its ownership is expressed.

A file states its own mode, and a whole tree can be forced one way from the command line: hecc --compatibility treats everything as legacy C for the initial drop-in, and you peel it back as you go. This is the -std= and TypeScript-strict pattern, not a new invention.

Two honest notes. Strict raises the floor -- no silent escapes -- but it is still not a guarantee: an origin can be stated wrongly, and the explicit untracked escape is always there, the way Rust's unsafe is. And strict is only as sound as its boundaries: a strict function calling into compatibility code is checked only as far as that code's contracts reach, which is exactly why annotated library headers matter.

Losing Ownership Information

There will be operations where hecc cannot reliably continue tracking ownership. For example:

int* pointer = malloc(...);

uintptr_t address = (uintptr_t)pointer;

followed later by:

int* other = (int*)address;

may destroy enough information that the compiler can no longer reliably determine the intended ownership relationship. The same applies to arbitrary pointer-to-pointer manipulation, indirect memory modification, complicated casts, and sufficiently unusual uses of memcpy(). hecc should not forbid these operations. Instead, there needs to be a clear boundary between code the ownership system understands and code where the programmer has performed an operation that makes reliable analysis impossible. At that point, hecc should simply stop pretending that it knows what happened. This is an important part of the project's philosophy. The compiler does not need to solve every possible C program. If the programmer gives it enough information, it can perform useful analysis. If the programmer deliberately removes that information through low-level operations, ordinary C remains available.

Bounds Analysis

Ownership is not the only area where additional information can improve C's safety. Consider:

int array[5];
int* pointer = array;

pointer[6] = 5;

The compiler knows that array contains five elements and that pointer originated from it. It can therefore detect that the access is outside the known bounds. This does not require changing the pointer into a special runtime type. The information already exists in the program.

It helps to think of bounds analysis in terms of how much the compiler knows, which tends to fall into three levels. A pointer can be completely unknown, where the compiler has no idea what it refers to. It can be inferable, where the pointer's origin gives the compiler bounds information on its own, as with pointer above. Or its bounds can be explicitly declared by the programmer, for the cases where the compiler cannot work them out but the programmer knows them.

When the information is already inferable, nothing extra is needed. When the programmer wants to state it directly, hecc can accept an explicit form:

int* pointer = (int[5])array;

This tells hecc to treat pointer as referring to an array of five int values for analysis purposes. It does not change the runtime representation of the pointer and does not need to exist in the generated C. The pointer is still an ordinary int* at runtime; the array information is there purely so the compiler can reason about it.

The same form covers the case where the compiler cannot infer anything at all:

int* pointer = get_memory();

The compiler does not know what this pointer refers to. It might be a single integer, an array, or video memory. The compiler should not pretend that every pointer has known bounds. Where the programmer knows the bounds but the compiler does not, they can be declared explicitly:

int* pointer = (int[BUFFER_SIZE])get_memory();

The syntax is still experimental, but the idea is to tell the compiler that the returned memory should be treated as an array containing BUFFER_SIZE integers for analysis purposes while remaining an ordinary pointer at runtime.

hecc should not prevent ordinary low-level C programming. This remains valid:

volatile uint8_t* vram = (volatile uint8_t*)0xB8000;

The compiler cannot meaningfully determine ownership or complete bounds information for arbitrary hardware addresses, and there is no reason to force them into the ownership system. This is not a failure of hecc's analysis. The information simply does not exist. The ownership system is a tool available when it is useful. It should not become a requirement for every pointer in every program.

Array parameters

The same bounds information can be attached to function parameters, and here C already has the syntax. It just throws the information away. Three forms are worth separating:

void foo(int array[]);
void fizz(int array[5]);
void buzz(int array[n], size_t n);

All three decay to int* at the ABI level, exactly as they do today, so nothing about the generated code or the calling convention changes. What differs is how much hecc knows.

int array[] says nothing. It is int* with a suggestive spelling, and hecc treats it as an unknown pointer. This is the escape hatch and stays the default.

int array[5] is a compile-time bound. Inside fizz, an access at index five or beyond is a provable violation, and a call passing an array hecc knows to be smaller can be flagged. This is the same idea as C's existing int array[static 5], which already means "at least five elements, and not null." hecc leans on that form rather than inventing a parallel one, and extends the reasoning it enables.

int array[n] is a runtime bound tied to a companion parameter. hecc carries the fact that array holds n elements, so an index into array wants to be below n, and the check connects to the loop analysis below. This is a fat pointer without the fat pointer: the length rides along in a value that is already being passed.

The bound may name a parameter that appears later in the list, as in buzz above. Standard C does not allow this. The name n is not in scope yet when array[n] is parsed, so the natural buffer, size order is an error, and the standard forces the reversed void buzz(size_t n, int array[n]). C23 added int array[.n] to forward-reference a later parameter, but few will reach for it.

This ordering rule is the whole reason the feature is adoptable. C code has written (buffer, size) for forty years. An existing declaration

void f(int* buffer, size_t size);

becomes

void f(int buffer[size], size_t size);

with one parameter's type changed and everything else left alone. Same order, same names, same ABI. Callers do not move and do not break. It is a pure annotation swap that gains the length fact for free. Requiring the size first would mean reordering the signature and touching every call site, which is exactly the friction that leaves a feature technically present and practically dead.

When the bound resolves to nothing, a misspelled or out-of-scope name, hecc does not treat it as a hard error the way standard C would. It falls back to [] semantics, an unknown length, and warns. This keeps the feature additive and non-fatal, in line with the rest of the analysis.

The length travels only as far as the annotations do. Assign array to a plain int*, or hand it to a [] parameter, and the fact is gone, the same way ownership is lost when it escapes into untracked code. One thing not to read into the syntax: int array[5] as a parameter is still a decayed pointer, so sizeof(array) inside the function is sizeof(int*), as in ordinary C. The bracket is a contract for analysis, not a real array passed by value.

Detecting Leaks

Ownership information can also help the compiler detect leaks. Consider:

int function(void)
{
    int* a = malloc(...);

    if (!a)
        return -1;

    int* b = malloc(...);

    if (!b)
        return -1;

    return 0;
}

If the second allocation fails, a is leaked. When hecc knows that a owns an allocation, it can follow the control-flow paths and detect that the function exits without releasing or transferring that ownership. The goal is not necessarily automatic cleanup. The generated program should remain close to ordinary C. Diagnostics are more useful here. The programmer remains responsible for freeing memory, but the compiler can point out when an ownership path appears to end without the allocation being released or transferred.

defer

The leak above happens because the cleanup is far from the allocation and easy to miss on one path. defer puts the cleanup next to the acquisition and runs it when the block exits:

int* buffer = malloc(...);
defer free(buffer);

/* use buffer, return early, whatever */

Whatever happens after -- an early return, a break, or falling off the end of the block -- free(buffer) runs first. The cleanup is written once, beside the thing it releases.

Deferred statements run in reverse order, as a stack:

acquire_a();
defer release_a();

acquire_b();
defer release_b();

acquire_c();
defer release_c();

At the end of the block the releases run release_c(), then release_b(), then release_a(). Reverse order is what you want when later resources were set up using earlier ones, so each is torn down before the thing it depended on.

A defer is scoped to its enclosing block, not the whole function, which is the choice Zig, Swift, and the C2y defer proposal make. A defer inside an if or a loop body runs at the end of that block, and in a loop it runs once per iteration.

Lowering

defer is a source transformation. The compiler inserts each deferred statement, in reverse order, before every exit from the block it belongs to, producing the C a careful programmer would write by hand:

int* buffer = malloc(...);

if (error)
{
    free(buffer);
    return -1;
}

free(buffer);
return 0;

It adds no runtime machinery and no hidden state, and code that does not use it is unchanged. The case that needs care is jumping out of a scope with goto or longjmp, where the end of the block is ambiguous; the --simple-control-flow ruleset restricts those and keeps the expansion well-defined.

defer and ownership

defer and the ownership system work together. A defer free(p) beside p's allocation is the discharge the leak analysis looks for, so an owned pointer with a matching deferred free needs no further proof that it is released.

It also catches a bug. If ownership of p leaves the scope before the block ends -- returned, or moved with <- -- the deferred free(p) would free something the caller now owns. hecc knows p no longer owns its allocation at that point, so it flags the deferred free rather than letting it run. hecc checks deferred cleanup against moved ownership.

Structs and Compound Types

Ownership extends to structures that contain pointers. Ordinary assignment is still an ordinary C copy:

b = a;

while an explicit transfer recursively moves the ownership information of the structure's members:

b <- a;

Whether a given pointer field is owned or borrowed is the distinction drawn earlier under Borrows in structures: an owned field is released with the structure, a borrowed field ties it to an origin the instance must not outlive. Nested, self-referential, and mixed-ownership structures still need careful rules, but those two field kinds are the foundation.

free()

Freeing memory should stay as permissive as it is in ordinary C. hecc is not trying to police every call to free(), and ordinary patterns must remain legal:

int* a = malloc(...);
int* b = a;

free(b);

This is allowed. The programmer has aliased the allocation and freed it through the alias, which is a perfectly normal thing to do in C. hecc should not reject it simply because the allocation was not freed through the pointer that originally owned it.

What hecc can do is warn. When memory is freed through a pointer that did not directly originate from a source the compiler understands, such as malloc(), a function known to return owned memory, or another explicitly understood source, hecc can point out that it cannot confidently account for what is being freed. The distinction is not between allowed and forbidden C. It is between code hecc understands with confidence and code where the programmer has deliberately given it less information. In the first case the compiler can reason about the free. In the second it still lets the program compile, but it may note that it can no longer be sure the operation is correct.

realloc()

realloc() is particularly complicated because it can resize an allocation in place, move it, or fail while leaving the original allocation valid. Ordinary C already requires careful handling:

int* temporary = realloc(pointer, size);

if (temporary)
    pointer = temporary;

Because realloc() participates in ownership, hecc can accept the explicit form:

int* temporary = realloc(<- pointer, size);

The <- communicates that the call takes part in ownership transfer, and the compiler applies what it knows about realloc() specifically. This is not an ordinary, unconditional move. realloc() is conditional. If it succeeds, it consumes ownership of the original allocation and the returned pointer receives ownership. If it fails and returns NULL, nothing has moved: the original pointer is still valid and still owns its allocation. hecc can encode both outcomes rather than treating the call as a single unconditional transfer.

This is a good example of the compiler understanding the special ownership semantics of a known function instead of forcing every operation through a completely generic rule. realloc() behaves the way it does because of how the C library defines it, and hecc can model that directly rather than pretending it is an ordinary assignment.

Type Unions

C frequently represents operations that can produce multiple possible outcomes through conventions rather than through the type system. A function may return NULL on failure, return a special integer value, modify an output parameter, set errno, or require the caller to inspect some other state. When a programmer wants to explicitly represent a value-or-error relationship, the usual C solution is to create a wrapper structure:

struct DivisionResult
{
    int result;
    enum DivisionError error;
};

This works, but it becomes repetitive when used throughout a larger API. Every possible return type requires another wrapper:

struct FileResult
{
    struct File file;
    enum FileIOError error;
};

struct StringResult
{
    char* string;
    enum FileIOError error;
};

struct SizeResult
{
    size_t size;
    enum FileIOError error;
};

These structures also represent two pieces of information simultaneously, even when they are logically mutually exclusive. A successful operation produces a value. A failed operation produces an error. The result is not necessarily an object that meaningfully contains both.

hecc introduces type unions to represent values that may contain one of several explicitly specified types:

int | enum DivisionError div(int a, int b)
{
    if (b == 0)
        return DivideByZero;

    return a / b;
}

The return value is either an int or an enum DivisionError. The function signature communicates this directly without requiring a separate wrapper type.

Type unions may contain more than two types:

int | char* | enum Error value;

A type union represents exactly one of its member types at a time. It is not equivalent to a structure containing every member simultaneously.

Type unions are also the first hecc feature that is not purely an analysis annotation. Ownership and bounds both erase completely at code generation: they add information for the compiler and cost nothing at runtime. A type union does not erase. It has a real runtime representation, a discriminator, a size, and an ABI. It is the one construct in hecc that spends runtime rather than only compiler attention, and that cost is deliberate.

The mechanism, however, stays small. Type unions introduce no new keyword. The only new syntax is the | type combinator, and existing C constructs such as typeof() are given a tighter definition rather than replaced. C programmers have wanted tagged unions for a long time, and this provides them without a separate ecosystem or a new keyword. The feature is large in consequence but minimal in mechanism, which is the trade being made on purpose.

Enums

A type union discriminates its members by type. That is enough for a value-or-error result, where the cases already differ in type, but it breaks down when two cases share a type: uint32_t | uint32_t is degenerate, so an event that is either a key-down code or a key-up code, both uint32_t, cannot be told apart. What separates those cases is not their type but their name.

The general form gives each case a name and an optional payload. This is C's enum with the one useful restriction lifted -- a variant may carry data -- and it turns out to be a single construct that covers three:

enum Shape {
    Circle(float radius),
    Rect(float w, float h),
    Empty,
};
  • A variant with no payload is an ordinary enumerator, so a plain enum Color { Red, Green, Blue } is the degenerate case: now a closed, distinct type the compiler checks exhaustively, which is the enum C should have had.
  • A set of variants told apart by type is the anonymous T | E type union, a shorthand for when the types alone distinguish the cases.
  • Variants with names and payloads are the general sum type, the thing Rust's enums are.

A value is built by naming the variant, and consumed by matching, which binds the payload to ordinary locals:

enum Shape s = Circle(2.0f);

switch (s)
{
case Circle(r):    return 3.14159f * r * r;
case Rect(w, h):   return w * h;
case Empty:        return 0;
}

This is the type-switch again, matching a variant name rather than a type: switch (typeof(x)) matches the cases of an anonymous union, switch (x) matches the named variants of an enum. The match must cover every variant or carry a default, the same exhaustiveness described below, and here it is simpler because it compares tags rather than resolving types. Inside an arm the bound names (r, w, h) have the variant's declared types; a bind borrows by default and moves only when the arm consumes an owned payload.

Because the variants are a closed, named set, this also fixes the rest of what makes C enums weak. The type is distinct -- an enum Shape is not an int and not another enum, and crossing to the underlying integer needs an explicit cast, which is the escape hatch and where a safe build can trap on an out-of-range value. The variant names are scoped to the enum, written Shape.Circle where a bare name would be ambiguous, so two enums no longer collide over a shared name, while the unscoped names stay available for existing code. And the tag's underlying integer type can be fixed, as C23 already allows: enum Shape : uint8_t.

Type unions and enums are therefore the same construct with two spellings. T | E is the anonymous form, its variants named by their type; the named enum form is the general one. Everything the following sections say about representation, discriminators, exhaustiveness, and ABI applies to both.

Existing C APIs

One of the main reasons for making hecc an extension of C rather than designing an entirely new systems language is the ability to gradually improve existing codebases. A project should not need to be rewritten in Rust, Zig, or another language before additional information can be communicated to the compiler.

Type unions are particularly useful for this purpose. Existing APIs can frequently be given more explicit result types without requiring a separate wrapper structure for every function:

FILE* | enum FileIOError open_file(const char* path);

char* | enum FileIOError read_file(FILE* file);

size_t | enum FileIOError write_file(
    FILE* file,
    const void* buffer,
    size_t size
);

The intention is not to require an entirely new hecc-specific standard library. Existing C libraries can continue to exist and ordinary C conventions can continue to be used where necessary. Type unions provide an additional way to describe the possible outcomes of an operation when the programmer wants the compiler to understand them.

This allows hecc features to be gradually introduced into an existing codebase. A project does not need to replace every API or rewrite every library before it can begin benefiting from additional static analysis.

One piece is deliberately unfinished. Representing a value-or-error is already solved: T | enum Error carries the outcome and, on failure, the reason. A plain value-or-nothing works the same way with a dedicated empty member, T | void, distinguished by the discriminator; there is no in-band sentinel. Returning NULL from a T | void* would collapse every failure to a single "nothing", discard the error reason, and collide with pointer types whose null value is a legitimate result, which is the in-band sentinel tagged unions exist to remove. Propagation is the open part. Once these types are returned, a caller may want a concise way to pass an error outward instead of type-switching at every level, like ? or try. Whether hecc grows such an operator or stays without one, leaving callers to check the type explicitly as C and Go do, is still undecided.

Runtime Representation

Unlike an ordinary C union, a hecc type union must know which member is currently active. This is necessary for runtime operations such as:

typeof(value)

An ordinary C union contains sufficient storage for all of its possible members, but it does not retain information about which member was most recently stored.

hecc type unions therefore use a discriminator.

Conceptually, a type:

int | enum DivisionError

is represented as:

enum DivisionResultType
{
    DivisionResultInt,
    DivisionResultDivisionError
};

struct DivisionResult
{
    enum DivisionResultType type;

    union
    {
        int integer;
        enum DivisionError error;
    } value;
};

The exact generated names are implementation details, but the representation itself is part of the language's ABI.

Every type union consists conceptually of two components:

  1. A discriminator identifying the currently active type.
  2. Storage large enough and correctly aligned for the largest member.

The discriminator and payload form a predictable, stable representation rather than compiler-private metadata. A type union does not require a garbage collector, runtime type system, or hidden runtime library. It is fundamentally equivalent to a conventional tagged union.

A named enum lowers the same way, with the union members named by variant rather than by type. This is what lets two variants share a payload type without colliding:

enum Shape {
    Circle(float radius),
    Rect(float w, float h),
    Empty,
};

becomes:

enum Shape_tag { SHAPE_CIRCLE, SHAPE_RECT, SHAPE_EMPTY };

struct Shape
{
    enum Shape_tag tag;

    union
    {
        struct { float radius; } circle;
        struct { float w, h; } rect;
        /* Empty carries no payload */
    } payload;
};

A variant with no payload contributes only a tag, so an enum whose variants all lack payloads has an empty union and lowers to a plain C enum -- an int at runtime, with the closed-set and exhaustiveness checking done entirely at compile time. That is why a C-style enum stays zero-cost while a payloaded one pays for its tag exactly as a type union does.

Discriminator Values

Discriminators are an internal encoding. The programmer never writes a discriminator value directly -- source code compares types, and the compiler lowers each type to an integer index for storage and dispatch (see Checking the Active Type). The encoding is described here because it is part of the ABI, not because it appears in ordinary hecc code.

The order of types within a type union determines those values.

For:

int | enum DivisionError

the conceptual mapping is:

0 = int
1 = enum DivisionError

For:

FILE* | enum FileIOError | int

the mapping is:

0 = FILE*
1 = enum FileIOError
2 = int

The exact discriminator representation should be defined as part of the hecc ABI. The implementation should use a predictable representation rather than allowing the backend to arbitrarily choose one.

This means that the ordering of types is semantically meaningful for ABI purposes. Reordering:

int | enum Error

into:

enum Error | int

changes the discriminator mapping and may therefore change the ABI of values exposed across compilation boundaries.

Checking the Active Type

The active member of a type union can be inspected using typeof():

int | enum DivisionError result = div(5, 0);

switch (typeof(result))
{
case int:
{
    auto value = (int)result;
    printf("%d\n", value);
    break;
}

case enum DivisionError:
{
    auto error = (enum DivisionError)result;
    printf("error: %d\n", error);
    break;
}
}

Each arm is its own scope, so it is braced when it declares a variable, the same as an ordinary C switch. Here typeof(result) yields the value's active type as a first-class type-value, and the switch matches it against types rather than integers. This is a type switch, not C's integer switch: case int: is a type, not an integer constant. It lowers to an ordinary integer switch on the internal discriminator, but at the source level types are compared directly. typeof() applied to a value gives that value's dynamic type; applied to a type it is the identity, so typeof(int) is simply int, which is why the labels can be written as bare types.

Checking the active type does not automatically change the static type of the variable. result remains:

int | enum DivisionError

even within the corresponding control-flow branch.

The programmer explicitly extracts a member through a cast, and auto can infer the member's type so the name is not written twice:

auto value = (int)result;

or:

auto error = (enum DivisionError)result;

This preserves the distinction between the union type and its possible members. A type union is not automatically treated as one of its members simply because the compiler can infer which member is active on a particular path.

Extraction and Construction

Extraction reuses C's cast syntax, but with one rule that removes an ambiguity. For a union whose members are not interconvertible, such as FILE* | enum FileIOError, a cast can only mean extraction. For a union whose members are interconvertible, such as int | float, an ordinary reading of (int)value is ambiguous: it could extract the int member, or it could numerically convert whatever member is active. Those produce different results:

int | float x = 3.5f;   /* float is active */
int y = (int)x;         /* extract int, or convert 3.5 to 3? */

hecc resolves this by rule: a cast whose operand is a type union is always extraction; a cast whose operand is a scalar is an ordinary conversion. To convert rather than extract, extract first and then convert:

int y = (int)((float)x);   /* extract the float, then convert it to int */

The inner cast targets a union, so it extracts the float. The outer cast targets a plain float, so it converts. There is no ambiguity because the two casts operate on different kinds of operand.

Construction works in the opposite direction. A value flows into a union when the compiler can determine which member the expression corresponds to:

int | enum DivisionError result = 5;   /* the int member */

When the member is ambiguous, the assignment is an error and the programmer states the member with a cast:

int | long n = 5;          /* error: int or long? */
int | long n = (int)5;     /* the int member */

A value constructs the member whose type matches it, and the ambiguity only arises when members are mutually convertible, as int and long are. An enum is treated as its own type here rather than a plain integer, so int | enum DivisionError has no such overlap and the div example above constructs each member from a bare return without a cast. This is the same cast syntax used for extraction, now selecting which member a value is stored as. No new keyword is required.

Exhaustiveness

A type switch must handle every member of the union, or provide a default. An incomplete switch is an error:

switch (typeof(result))   /* result : int | enum DivisionError */
{
case int:
    ...
    /* error: enum DivisionError is not handled */
}

This is also what makes extraction inside a switch safe without a separate check. Within each arm exactly one member is statically known to be active, so the cast in that arm -- (int)result under case int: -- is a validated extraction by construction. An if (typeof(result) == int) guard establishes the same fact on its true branch, and the compiler validates the cast there the same way. Checking the active type does not, on its own, change the static type of the variable -- result stays int | enum DivisionError, and the programmer still performs the cast -- so what the check buys is a proven-valid extraction, not automatic narrowing. The strict, warning, and permissive rulesets described next apply to extraction performed outside such a proven context, where the active member cannot always be established. Inside an exhaustive switch, or a guarded branch, it always can.

Unchecked Extraction and Compiler Leniency

There are situations where a programmer may deliberately want to extract a member without validating the active type first:

FILE* | enum FileIOError result = open_file(...);

FILE* file = (FILE*)result;

The cast explicitly communicates that the programmer expects result to currently contain a FILE*.

Whether this operation is accepted depends on the selected compiler strictness or ruleset.

A strict configuration may reject the operation:

error: cannot extract FILE* from FILE* | enum FileIOError
without validating the active type

A less strict configuration may allow it while producing a warning:

warning: unchecked extraction of FILE*
from FILE* | enum FileIOError

A permissive configuration may accept the cast without diagnostics.

This allows type unions to be gradually introduced into existing C codebases. A project can initially use permissive rules while adding additional type information, then progressively enable stricter checking as the codebase becomes more explicit.

The cast therefore acts as an explicit unchecked assertion by the programmer. The compiler may reject or warn about the assertion depending on the selected ruleset, but the language does not need to make unchecked operations universally impossible.

This is a force unwrap, and it has two independent parts. At compile time the compiler rejects, warns, or accepts it, which is the strict, warning, and permissive spectrum above. At runtime an accepted unchecked cast does nothing extra by default: it erases to a plain member access, costs nothing, and reading the wrong member is undefined behaviour, as in ordinary C. That keeps with hecc's rule that extensions do not silently inject runtime behaviour the programmer did not write.

The T | void case makes the hazard concrete. The void member has no payload, so if result currently holds the empty case, a force unwrap:

int | void result = some_function();
int number = (int)result;

reads storage that was never written. Under the default rules this is unwrap_unchecked, not a checked unwrap.

Because the discriminator exists at runtime, a checked unwrap is nearly free to provide, and it fits the opt-in rulesets rather than the base language. A rule such as --checked-unwrap would turn every force unwrap into a tag comparison followed by a trap when the asserted member is not active. In a C backend the trap is an abort -- abort(), __builtin_trap(), or a user-supplied handler -- not an unwinding panic, because hecc has no unwinding runtime. This gives a debug-and-release split for free: build with the rule and a wrong force unwrap aborts at its source; build without it and the same casts erase to nothing. It should not be the default, because trapping is runtime behaviour the programmer did not ask for, but it is a cheap thing to offer precisely because the tag is already there -- something C cannot do at all.

Stable ABI

The representation of a hecc type union is part of the language ABI.

This is particularly important because hecc is intended to coexist with existing C code rather than create an isolated ecosystem. A type union exposed through a library boundary should have a representation that can be understood and reproduced without requiring another hecc compiler.

For example, a C implementation should be able to construct and consume a value corresponding to:

int | enum DivisionError

by following the specified ABI using ordinary C constructs:

enum DivisionResultType
{
    DivisionResultInt,
    DivisionResultDivisionError
};

struct DivisionResult
{
    enum DivisionResultType type;

    union
    {
        int integer;
        enum DivisionError error;
    } value;
};

A function implemented in C could therefore return the equivalent representation:

struct DivisionResult div(int a, int b)
{
    if (b == 0)
    {
        return (struct DivisionResult)
        {
            .type = DivisionResultDivisionError,
            .value.error = DivideByZero
        };
    }

    return (struct DivisionResult)
    {
        .type = DivisionResultInt,
        .value.integer = a / b
    };
}

Likewise, a hecc program can consume values produced by C code that follows the same representation.

This makes type unions suitable for:

  1. C bindings.
  2. Foreign function interfaces.
  3. Shared libraries.
  4. Mixed C and hecc projects.
  5. Assembly interfaces.
  6. Gradual migration of existing codebases.

The exact in-memory layout should be formally specified, including discriminator representation, payload alignment, padding, and member ordering. Calling conventions do not need to be invented separately. A type union can follow the target platform's existing ABI rules for passing and returning the equivalent C structure.

This distinction is important. hecc defines what the type looks like in memory, while the platform ABI determines how an equivalent aggregate is passed between functions.

This also settles how a type union crosses a header. A hecc consumer can include a header written in hecc syntax and see int | enum DivisionError directly. A plain C consumer uses the lowered form instead -- the transpiler can emit a corresponding C header declaring the generated tagged struct -- and reads the result through its type and value members like any other tagged union. Both sides agree because they share the specified ABI. The convention still to pin down is which header is canonical in a mixed project: whether hecc emits the C-facing header from the hecc one, or the two are maintained side by side. A library written in hecc is therefore still callable from C through its lowered struct. What C cannot do is parse the hecc-syntax header, which is expected.

Transpiling Type Unions

The stable representation also makes type unions straightforward to support in a C backend.

A hecc function returning int | enum DivisionError:

int | enum DivisionError div(int a, int b)
{
    if (b == 0)
        return DivideByZero;

    return a / b;
}

lowers to a function returning the tagged union shown earlier -- the DivisionResult struct with its discriminator and payload -- and each return becomes an assignment of the corresponding member and tag. The generated C contains no magical runtime machinery. The extended hecc type becomes ordinary C constructs that represent exactly the semantics defined by the language.

A named enum lowers by the same rule, with the only new codegen being the payload binding in a match. Construction is a designated initializer:

enum Shape s = Circle(2.0f);
struct Shape s = { .tag = SHAPE_CIRCLE, .payload.circle = { .radius = 2.0f } };

and a match becomes a switch on the tag, with each bound name declared as a local drawn from the active union member:

switch (s)
{
case Circle(r):    return 3.14159f * r * r;
case Rect(w, h):   return w * h;
case Empty:        return 0;
}
switch (s.tag)
{
case SHAPE_CIRCLE: { float r = s.payload.circle.radius;                    return 3.14159f * r * r; }
case SHAPE_RECT:   { float w = s.payload.rect.w, h = s.payload.rect.h;     return w * h; }
case SHAPE_EMPTY:  {                                                       return 0; }
}

The binding is the whole of the new work: reading the fields of the active variant into locals. Everything else is the tagged-union lowering already described. A recursive enum, one whose variant refers to itself, takes a pointer in that variant exactly as a recursive C struct does, and ownership of that pointer follows the owned-field rules; that is the one case that needs the care structures already need.

An LLVM backend can use the same logical representation directly while preserving the ABI rules defined by hecc.

Why Type Unions Exist

C can already build tagged unions by hand from structs, unions, and enums; the problem is repetition, since every new value-and-error combination needs another wrapper with its own discriminator, payload, naming, and extraction logic. A type union expresses the relationship directly and lets the compiler provide that machinery while the programmer keeps a predictable, explicitly defined representation. This fits the broader purpose of hecc: add information where C currently relies on convention, boilerplate, or undocumented discipline, in a way that stays interoperable with ordinary C through a stable ABI.

Because the discriminator exists at runtime, a type union can also serve as a runtime-polymorphic value: a struct Circle | struct Rectangle | struct Triangle inspected with a typeof switch gives single dispatch -- polymorphism without inheritance, vtables, or templates -- over a closed, known set of types. This is a side effect of the representation rather than a goal, and it does not replace generics, which work for an arbitrary type chosen at compile time. Where a program only has to handle a fixed set of runtime types, listing them is simpler than a generic abstraction; where it needs an arbitrary type, a type union is the wrong tool.

Nullability

There is one union that needs no tag at all, and it is the most common one in C: a pointer that may or may not point at something. T* today conflates "points at a T" with "might be null," and that conflation is behind a large share of C's crashes. hecc lets the pointer say which it is:

int *nonnull  p;   // never null
int *nullable q;   // might be null
int *         r;   // unknown; today's behaviour, and the default

An unannotated pointer keeps exactly its current meaning, so existing code is unchanged and the qualifiers are something you opt into. All three are bit-identical to an ordinary int* at runtime. There is no tag and no extra storage, because a pointer already carries its own empty case: the null value is 0, which is already a valid pointer pattern. nonnull and nullable only state whether that value is allowed. This is why the earlier point about int* | void does not apply here. A type union is tagged and would make the pointer fat and change its ABI; nullability instead refines the inhabitants the pointer already has, so it stays a plain int*. It is the union that costs nothing because one of its arms is a value the carrier was already able to hold.

The qualifier sits after the *, in the same slot const uses when it describes the pointer rather than the pointee:

int *const    p;   // p is a const pointer to int
int *nullable p;   // p is a nullable pointer to int

That placement matters when several pointers are declared together. After the *, each qualifier belongs to its own declarator and does not leak onto the next:

int *nullable p, *nonnull q;   // p is nullable, q is nonnull

Written before the * the qualifier would sit with int in the declaration specifiers, which apply to the whole declaration, and it would smear onto every declarator in the list. After the * avoids that, the same way int *const a, *b; leaves b an ordinary pointer.

What the checker does with this is refuse to dereference a nullable pointer until it has been tested. The test narrows the type, exactly as a type-switch narrows a union to its active arm:

void use(int *nullable p)
{
    *p = 5;        // flagged: p may be null

    if (p) {
        *p = 5;    // fine: p is nonnull in this branch
    }
}

A nonnull parameter moves the check to the boundary. The caller has to establish that the pointer is not null before the call, and inside the function it can be used directly, which pushes the one test to where the information actually is instead of repeating a defensive if in every callee.

This is the same state tracking the ownership system does -- a value sits in one of a few compile-time states and an operation is legal only in some of them -- narrowed here to the two-state case of a pointer. Ownership (owned, moved, freed) and nullability (null, non-null) are the same mechanism with different states; the general, user-defined form is typestate, in the explorations.

Transpilation

The initial implementation of hecc is likely to be a transpiler. The compiler would parse the extended syntax, perform semantic and ownership analysis, and generate ordinary C:

hecc source
    |
    v
parser and analysis
    |
    v
generated C
    |
    v
C compiler
    |
    v
binary

Many of the extensions can disappear completely after analysis.

b <- a;

becomes:

b = a;

and:

int* pointer = (int[5])array;

becomes:

int* pointer = array;

The intention is to keep the generated C as close to the original program as possible. hecc should primarily add information for analysis rather than secretly generate a fundamentally different program.

This also gives the project two independent exit doors, which matters for a language that is asking people to adopt a dialect. A project can transpile once, commit the generated C, and continue as an ordinary C project that no longer needs hecc at all. Or it can keep hecc as its compiler and treat the generated C as a fallback it can drop back to at any time. Either way the source is never trapped inside the hecc toolchain: if the compiler stops being maintained, the code it has already produced does not stop working. This is the main structural difference from earlier safe-C dialects that relied on a runtime and could not hand back clean C.

A future LLVM backend could use the same frontend:

              hecc source
                   |
                   v
           parser and analysis
              /           \
             v             v
        C backend      LLVM backend

The C backend is the simplest place to begin because it allows experimentation with the language without requiring an entire backend immediately.

The same code-generation step can occasionally make the emitted C faster rather than only safer, by turning facts the analysis proved into hints the backend understands: a non-aliasing it established becomes restrict, a proven bound becomes a __builtin_assume that lets the backend drop a redundant check, a nonnull becomes the platform's non-null attribute. hecc only emits a hint it can back, never a guessed one, since a wrong restrict is undefined behaviour rather than a missed warning. It is a minor, opt-in payoff and not a goal, but it is unusual for a safety layer to hand any cycles back at all.

Implementation

Transpiling to C keeps the compiler small. Everything downstream of the emitted C -- register allocation, instruction selection, code generation, optimization -- belongs to the existing C compiler, so hecc is a frontend, an analysis stage, and a pretty-printer, and little else. Most of what makes a compiler large is inherited rather than written.

The analysis is where the work is, and two decisions already made in the language design happen to make it cheap to build.

The first is the diagnostics-not-guarantees stance. Because hecc warns only on what it can prove wrong rather than proving the absence of error, the analysis is allowed to be approximate: it can cap iteration, give up on a case it cannot follow, and accept false negatives, since a missed warning is within its remit. A sound checker cannot do any of that -- it must reach a precise fixpoint and account for every path, which is a large part of why Rust's borrow analysis, and Polonius in particular, are expensive. The same polarity that makes hecc quieter to use makes its analysis far less costly to compute. Precision becomes a dial, turned only as far as the diagnostics stay useful.

The second is that ownership contracts live on signatures. A function is checked against the summaries of the functions it calls -- which arguments are consumed, which returns are owned, which are borrows -- and never needs to see their bodies. That makes the analysis intraprocedural, roughly linear in the size of each function, and parallel across functions once the summaries are collected; whole-program analysis is never required. The annotations are not only for the programmer, they are what keeps the analysis from going quadratic.

Underneath, the substrate is ordinary compiler engineering. An arena holds the AST and IR and is released per translation unit, which is faster than per-node allocation and simpler, since there are no individual lifetimes to track -- the arena pattern hiv describes, turned on the compiler itself. Identifiers are interned to integers at lex time so every later comparison is a single word. The flow-sensitive checks -- ownership, leaks, borrows -- run as a standard forward dataflow over each function's control-flow graph: a lattice, transfer functions, and a worklist to a fixpoint, with bounds as interval analysis in the same frame. None of it is exotic; it is the textbook shape, kept cheap by the two decisions above.

Some of this is worth doing from the start because it is structure rather than optimization -- the arena, the interner, and the CFG-and-worklist form of the analysis are the correct shape, and retrofitting them later is a rewrite. The rest waits until a profiler asks for it: parallelism across functions, flat index-based node arrays in place of pointers, and incremental caching that memoizes a function's analysis on a hash of the function and the summaries it depends on. That last one is what would eventually give editor-grade responsiveness, so the passes are worth writing as pure functions of their inputs now, before anything is cached, to keep the door open.

The honest prior is that a transpiler at this scale will not be slow, and the real risk is over-engineering the compiler before the analysis is correct. The arena and the interner earn their place on day one for simplicity, not speed; everything past that waits for a measurement.

Runtime Checks

Everything above is static: the compiler proves what it can, warns about the rest, and it all erases. Some safety cannot be established statically, though -- particularly over the plain = code hecc deliberately leaves unchecked so it can be adopted at all. For that, hecc can insert runtime checks.

Because runtime checks cost cycles, they are never on by default; that would penalize adopted code and break the additive rule. They are opt-in build tiers, in the spirit of -fsanitize or Zig's safe release modes -- hecc --safe, or a granular selection like --checks=bounds,ownership,tags,overflow. A release build inserts nothing.

This layer pairs with the static one rather than duplicating it. Static analysis covers the parts written with <- and origins, for free, at compile time; runtime checks cover what static cannot see -- most importantly the unchecked = code -- so the two together reach far more than either alone, and a check the compiler already discharged statically can be elided.

The checks hecc is well placed to insert, because it already has the information for them:

  • Extraction (already described). A force unwrap (T)union becomes a tag comparison and a trap on mismatch. Nearly free, since the discriminator already exists.
  • Bounds. Where bounds analysis knows a pointer's extent -- inferred, or declared with (int[N]) -- an index or dereference past it traps. This is more targeted than a generic sanitizer's shadow memory: hecc checks against the intended extent it already knows, and inserts nothing where it has no bounds.
  • Ownership. Because hecc knows the owner and where free happens, a safe build can poison freed allocations, trap on a dereference of a borrow whose owner has been freed, and trap at free on a double free or a free through a non-owner.
  • Signed overflow. hecc can define signed overflow as two's-complement wraparound outright -- the behaviour most hardware already has, close to -fwrapv -- so it stops being undefined; or, in a safe build, trap on it instead. A defined result or a loud one, rather than C's silent UB. This is the one undefined behaviour hecc pins down by default; the approach is case by case, taking a single UB and giving it a predictable result where the cost is acceptable, not trying to fix all of C.
  • Null (optional). Trap on a dereference the analysis cannot prove non-null.

On a violation the program traps -- abort(), __builtin_trap(), or a user-supplied handler -- not an unwinding panic, since hecc has no unwinding runtime.

Two honest limits. First, these make a violation fail fast rather than corrupt silently; they do not make the program correct, only loud. Second, hecc's runtime layer is targeted at the classes it models -- bounds, ownership, tags, overflow -- not a general memory sanitizer; for arbitrary heap corruption or uninitialized reads a tool like AddressSanitizer is still broader, and nothing stops you running both. What hecc has is that its static information lets it insert cheaper, narrower checks than a tool working blind, and skip the ones it already proved.

Diagnostics

Because hecc tracks the flow of ownership, origins, and bounds, it knows why something is wrong, not just that it is, and the diagnostics should show that. A use-after-free is not a line number; it is a path -- here is where the allocation was owned, here is where ownership moved out, here is where the borrow was used afterward. hecc has all three points and should print all three, in the style of Rust's or Elm's errors rather than a terse C compiler's.

This is a design goal rather than a feature, but it is not a cosmetic one. The tools in this space that people actually keep using are the ones whose output tells you what to do next. An analysis that is right but illegible gets silenced, and a silenced checker checks nothing. For hecc specifically, the annotations are opt-in, so a diagnostic also has to justify the annotation that produced it -- to show that writing <- or an origin bought a real, comprehensible check -- or people will not bother writing them.

hiv

Ownership and strict mode are only as good as the contracts on the functions a program calls, and most of those calls land in the standard library. hecc understands a few standard functions directly -- malloc, free, realloc, strdup -- but the platform libc is otherwise an opaque boundary the analysis has to trust.

hiv is hecc's standard library, meant to close that boundary -- a reimplementation of the C standard library, itself written in hecc, so its functions carry ownership decorators natively: which arguments are consumed, which returns are owned, which are borrows and from where. That turns the whole standard library into an annotated boundary strict mode can check across instead of a wall the analysis goes dark at, and because hecc transpiles to ordinary C, hiv is still a normal C library for non-hecc callers.

hiv gives every allocating function two forms. The default keeps C's signature and carries ownership -- malloc(size), strdup(s), free(p) -- so existing code compiles unchanged while hecc still knows what is owned; the _with form takes an explicit allocator -- malloc_with(&a, size), free_with(&a, p) -- for when the caller wants to choose where memory comes from. The _sized variants pass along the buffer size the caller already knows, which feeds hecc's bounds analysis, and compose with _with.

hecc and hiv reinforce each other. hecc tracks who owns and who frees; hiv states the ownership of every call and makes the allocator reachable through _with. And an hiv guarding allocator -- poisoning freed memory, quarantining, guard pages, double-free detection -- is the runtime net from Runtime Checks, provided at the library boundary rather than by codegen. Static hecc over a guarding hiv allocator is the static-plus-runtime story end to end.

None of it is required. hecc works against the platform libc, and hiv is usable as an ordinary C library without hecc; it reimplements the ISO C interface rather than inventing a separate one, so it is not a new ecosystem to buy into. Pairing them is the maximal configuration, not a precondition.

Lexical Conveniences

A few conveniences live entirely in the lexer: they change how source is spelled, not what it means, and disappear before analysis or code generation. Numeric literals may use _ as a digit separator -- 1_000_000, 0xFF_FF -- which C's grammar already tokenizes as a single number and only rejects when converting it to a value, so hecc just relaxes that step rather than touching the tokenizer. Binary literals 0b1010, already standard in C23, and an explicit octal 0o755 in place of C's error-prone leading-zero 0755, round out the numeric forms. Block comments nest, so /* ... /* ... */ ... */ closes where you expect. Each of these strips to the plainest possible C -- 1000000, 10, 493, nothing -- and costs nothing.

Raw string literals are the one heavier member of the set, and they earn it. Spelled as C++ spells them -- R"(...)", with an R"tag(...)tag" form when the contents contain )" -- they take everything between the delimiters verbatim, newlines and backslashes included, with no escape processing. That is what makes an embedded shader, regex, or block of SQL bearable instead of a wall of \n\". Unlike the others this is not a pass-through: C has no raw strings, so hecc re-escapes the contents into an ordinary C string literal on the way out. It still erases to a plain const char[] and touches neither the analysis nor the ABI, so it stays lexer sugar -- it just earns its place by removing real pain rather than by being free.

The list stays deliberately short. Cosmetic syntax earns a place only when it either costs nothing or removes real, recurring pain, and only when it matches where C or the C family already leans. It is not where hecc spends its attention, and the bar for adding to it is high on purpose.

Keeping It Small

The biggest risk to hecc is the one every C successor has hit: solving each hard case by adding another keyword, type, or library component until the small language is a large one. The guard against it is a hard line between the core and everything else.

The core is what hecc claims to be, and it is deliberately small: ownership, borrows and origins, leak detection, and defer; bounds and array parameters; and sum types -- enums and the anonymous T | E unions -- with nullability. Ownership and nullability are the same idea underneath -- tracking which of a few compile-time states a value is in and gating operations on it -- so they count as one mechanism, not two. These are load-bearing and mutually reinforcing, they all erase to ordinary C, and none is required to compile existing code. That set is the language.

Even there, keywords are kept scarce. nonnull and nullable are contextual: each means something in exactly one position and reserves the word nowhere else, so existing code using those identifiers keeps working. defer is the one keyword that introduces a statement, and it earns that: it removes a whole class of leak, it lowers to ordinary C, and it matches the direction C itself is taking with the C2y proposal.

Everything past this point is the other side of the line: explorations held to the same discipline -- additive, erasable, opt-in -- but not yet part of what hecc is. Some may earn their way into the core, some may stay experiments or move into a separate tool, and some may be dropped. Keeping them visibly separate is what lets the core stay small while the thinking stays open. C's size is part of its appeal, and a large hecc-specific standard library or a pile of always-on features would work against the whole point. The compiler understands a few standard functions -- malloc(), free(), realloc(), strdup() -- without requiring a new standard library to be adopted first.

Beyond the Core

The rest of this document is exploration. Each piece holds to the same rules as the core -- it adds information, it erases to ordinary C, and it is opt-in -- but none is settled, and none is load-bearing for the features above. They are here because they are worth thinking through, not because they are decided. Typestate is the general form of the state tracking ownership and nullability already do; effects and distinct types are separate good ideas that would each stand on their own; and they add little to the surface -- the effect contracts and distinct are contextual keywords, and typestate adds none at all, reusing distinct, :, |, and ->.

Typestate

Nullability is a small instance of a larger pattern, and so is ownership. A nullable pointer is a value in one of two states -- null or non-null -- where a dereference is legal in one and not the other, and a test moves it between them. An owned pointer has the same shape: owned, moved-out, or freed, where free is legal on an owner and a bug on a pointer that has already been freed or moved. Both are the same mechanism running on fixed, built-in states: a value moves through a set of compile-time states, and an operation is allowed only in the states where it makes sense.

Typestate exposes that mechanism directly, so a programmer can declare the states themselves. The states are compile-time only; the value is still an ordinary FILE* or int at runtime, and the transitions erase like the rest of the analysis. A distinct type can list the states it moves through, and a function can say which state it requires and which it leaves behind:

typedef distinct FILE File : open | closed;

File* open_file(const char* path) : open;          /* result is in state open */
size_t read(File* f : open, void* buf, size_t n);  /* requires open, stays open */
void   close(File* f : open -> closed);            /* open on entry, closed after */

The : state on a parameter is a requirement; the a -> b is a transition. hecc tracks the state through the caller and flags an operation used in the wrong one:

File* f = open_file("data");

read(f, buf, n);   /* fine: f is open */
close(f);          /* f: open -> closed */
read(f, buf, n);   /* warning: read requires open, f is closed */

The interesting case is the one C has never been able to state: an API protocol, the order a library must be called in. "Open before read, do not use after close," "lock before touching the field, unlock after," "init before any other call" -- these are today enforced only by documentation and habit. Typestate lets the header carry the rule so the compiler checks every caller against it, exactly as ownership and nullability already do for their built-in states.

This is the most speculative feature here and the heaviest conceptually, so it stays firmly opt-in and firmly diagnostic. It proves nothing, forbids nothing, and an = or a cast drops the state the same way it drops ownership. What it buys is that the two features it generalises stop being special cases the compiler hard-codes and become the visible instances of one mechanism the programmer can reach for -- which is the trade hecc prefers: one more mechanism that removes the need for several, rather than several that each cover one case.

Effects

C has a couple of function attributes in this area -- pure and const -- but they are leaf annotations a programmer asserts and the compiler mostly trusts, not a system that checks and propagates. hecc can carry a small set of effect contracts and enforce them across the call graph: whether a function allocates, whether it can fail, whether it can block.

void  render(struct Scene* s)  no_alloc;   /* allocates nothing, transitively */
int   step(struct Machine* m)  no_fail;    /* has no failure path */
void  poll(struct Device* d)   no_block;   /* never blocks */

The contract is checked, not assumed. A no_alloc function may not call malloc, or any function that is not itself no_alloc -- and because hecc can inspect bodies it can see through to the leaf and infer the effect where it is not written down. What a programmer gets is the thing embedded, kernel, and real-time code always wants and C never offered: a way to prove a whole call path never allocates, or never blocks, checked at compile time rather than discovered in the field. It pairs directly with the bounded-loops ruleset below; both exist for worst-case-execution reasoning.

no_fail connects to type unions. A function that returns T | enum Error has its failure in the type, and no_fail is the claim that the error arm is never taken -- so the two describe the same property from opposite ends, one in the return type and one in the contract. Like everything else, effects are contextual keywords, opt-in, and erasable: they add no runtime code and vanish in the generated C, leaving only the diagnostic behind. An external function the compiler cannot see needs the contract on its header, the same as ownership.

Distinct Types

A typedef in C is a transparent alias. typedef int Celsius; gives Celsius a name, but it is still int in every way that matters: you can assign a raw int to it, pass a Fahrenheit where a Celsius is wanted, or hand a plain count to a function expecting a user id, and the compiler says nothing. The name documents intent that the type system does not enforce.

hecc adds a marker that makes the alias a type of its own:

typedef distinct int Celsius;
typedef distinct int Fahrenheit;
typedef distinct int UserId;

Celsius, Fahrenheit, and UserId are now three different types, distinct from int and from each other even though all three are represented as int. Mixing them, or moving between one of them and a raw int, needs an explicit cast:

Celsius c = 20;              // fine, an integer literal used as Celsius
Fahrenheit f = c;            // flagged: different types
int temperature = c;         // flagged: distinct type into plain int
Celsius c2 = (Celsius)68;    // explicit, allowed

This catches the boring, expensive class of bug where two quantities share a representation but not a meaning: temperatures in the wrong scale, an order id used as a user id, a byte count passed as an element count. The cast is the escape hatch, in the same spirit as = for ownership. When you genuinely mean to cross the boundary you say so, and the crossing is visible in the source.

It stays additive and erasable. A plain typedef behaves exactly as before, so distinct is opt-in, and a distinct type is its underlying type at runtime with the same size, layout, and ABI. It exists only for the compiler's checking and disappears in the generated C, where Celsius is once again just int. It earns a spot only if the unit- and id-mixups it catches are a real problem for a given codebase; it sits below the line because it is orthogonal to the memory and state safety that is hecc's actual subject.

Rulesets

Not every problem in C comes directly from memory management or the type system. A significant number of problems also come from the way programs are structured. C gives programmers a great deal of freedom in this area: control flow can jump arbitrarily through a function, and loops can run indefinitely.

hecc could provide optional, enforceable rules intended to encourage simpler and more predictable code. These rules are not part of the language itself, and none are enabled by default. A programmer should always be able to write ordinary C-like code without any of them. Projects opt into whichever rules suit the guarantees or style they want, and the rules could equally well live in a separate checking tool rather than the compiler. They are advisory conveniences for producing code that is easier to debug and analyse, not features of the language.

The general idea is that many problems can be avoided not by adding more abstractions to the language, but by reducing unnecessary complexity. Simpler control flow is easier to analyse, and bounded loops make execution easier to reason about.

1. Simple Control Flow

Unrestricted control flow can make programs significantly harder to reason about. C provides mechanisms such as goto and longjmp() which allow execution to move through a program in ways that are not represented by ordinary structured control flow. An optional rule could prohibit these constructs and require programs to use structured control flow instead.

This is not intended to suggest that constructs such as goto are universally useless. They are sometimes used deliberately, particularly in low-level C code. However, arbitrary jumps make control-flow analysis more complicated and can make ownership analysis, leak detection, and general reasoning about program state substantially harder. Restricting control flow to structured constructs gives both the programmer and compiler a clearer understanding of how execution can move through a function. This is particularly useful for hecc because ownership may move through different control-flow paths, and simpler control flow makes those paths easier to analyse. This behaviour could be enabled through a compiler flag: --simple-control-flow When enabled, constructs such as goto and longjmp() could be rejected or diagnosed by the compiler.

2. Bounded Loops

For code that needs worst-case execution guarantees -- embedded systems, real-time work, anything where unbounded iteration is a hazard -- it can be useful to require that every loop has a known maximum number of iterations. This is not a claim that every loop can be bounded, and it is not a default. It is an opt-in rule for projects that want it.

Some loops already carry their bound. For a simple for loop the compiler can see it directly:

for (size_t i = 0; i < 100; i++)
    ...

The loop cannot execute more than 100 times. Others do not, and under this rule the programmer supplies the bound. Consider a loop that polls a device until it is ready:

while (!device_ready())
    ...

Nothing here bounds the number of iterations. In a context that wants worst-case guarantees, the programmer caps it:

for (size_t i = 0; i < MAX_POLLS; i++)
    if (device_ready())
        break;

The loop can still exit early; the point is that there is an upper bound on how long execution can stay inside it. The exact syntax for declaring a bound on loops that cannot be inferred is still to explore. Bounded loops make worst-case execution easier to reason about and static analysis more predictable, which is why they are useful in the domains that want them -- and why they stay optional everywhere else. This behaviour could be enabled through: --bounded-loops. The compiler would enforce bounds where they can be determined and require an explicit bound where they cannot.

Prior Art and Answers

hecc is not the first attempt to add safety information to C, and the fair question is why it should fare any better than the attempts that did not survive.

Cyclone was the closest ancestor: a safe C dialect from AT&T and Cornell. It did not last, and part of the reason is that its safety leaned on runtime machinery such as fat pointers, so its output was not ordinary C. hecc avoids this deliberately. Ownership and bounds erase at code generation, and the transpiler emits plain C, so a project is never trapped inside the toolchain. If hecc stops being maintained, the C it has already produced keeps working. That is the one structural advantage over the dialects that came before, and it is the honest answer to "dialects die": the compiler may, the code does not.

Zig and Rust solve problems hecc does not attempt to, and they solve them better -- in Zig and in Rust. That is the point. Adopting either means porting; neither retrofits an existing C codebase in place. Zig already has tagged unions and error unions, more ergonomic than hecc's, inside Zig. Rust already has sound ownership, inside Rust, at the cost of a borrow checker and of the unrestricted aliasing C depends on. hecc is aimed at the code that is and will stay C, which those languages cannot reach without a rewrite.

The closest match in spirit is not another systems language but gradual typing. TypeScript, mypy, and Sorbet are unsound by design. They can be silenced with an escape hatch, adopted one file at a time, and they check intent the programmer states rather than intent the tool has to guess -- and they are enormously useful anyway. hecc applies the same bargain to memory responsibility. <- is closer to const than to a Rust move: an advisory contract the compiler checks where it can, defeatable on purpose, useful even though it is defeatable. Nobody argues const is worthless because a cast can strip it.

This framing also fixes the limits in place, honestly:

  1. hecc does not make programs memory-safe. It gives diagnostics, not guarantees. Double-free and use-after-free remain expressible.
  2. Ownership analysis is best-effort. Borrows and origins reconnect references that escape through returns, struct fields, and annotated signatures, but the checker warns rather than proves, external functions and pointer-to-integer casts still make it stop pretending, and nested or self-referential structures need more rules.
  3. Error propagation has no sugar, and may never. The value representation is settled; whether a ?-style operator is worth its weight is not.
  4. The rulesets are advisory, off by default, and could live in a separate tool. They are not part of the language.
  5. Typestate and effects are the most speculative parts and the least worked out. They generalise mechanisms hecc already needs -- nullability and ownership state, allocation and failure contracts -- but user-defined states, effect inference across large call graphs, and the exact syntax are all still open.

These are real. The claim hecc is willing to defend is narrow, and it is the one worth making: C with tagged unions and best-effort ownership and bounds diagnostics, adoptable gradually, interoperable through a stable ABI, with no path that traps a codebase inside the hecc toolchain.

Other C Dialects and Extensions

hecc is far from the first attempt to improve C. The others fall into a few distinct approaches, and hecc belongs to only one.

Evolution languages -- C2, C3. These clean up C itself: modules instead of the preprocessor, better error handling, slices, some generics. They are genuine improvements, but they are separate languages with their own compilers -- you write C2 or C3, you do not upgrade your existing .c files in place. hecc is the opposite bet: not a nicer language to switch to, but a layer added on top of the C you already have.

Gradual safety extensions -- Checked C, clang's -fbounds-safety, Sparse. These are hecc's closest relatives: annotations added to real C, checked incrementally, backward compatible, adoptable file by file. The difference is what they check. Checked C and -fbounds-safety add spatial safety -- checked pointer and array types that carry their bounds -- and the Linux kernel's Sparse checks address spaces and locking through __user, __rcu, and similar tags. All of them share hecc's additive, tool-checked model, and Sparse in particular is proof it scales: the kernel has run on exactly this kind of opt-in annotation for two decades. hecc overlaps them on bounds, but its main investment is on the axis they mostly leave alone -- temporal safety: ownership, use-after-free, double-free, leaks -- plus type unions. Where a project wants spatial safety, -fbounds-safety is real and shipping today; where it wants ownership, that is hecc's ground.

Restrict-and-verify -- MISRA C, Frama-C. MISRA restricts C to a safer subset enforced by linters; hecc's optional rulesets are a lighter version of the same idea. Frama-C goes the other way, adding a contract language for formal proof of properties. hecc sits between the two: more than a lint, far less than a proof -- best-effort diagnostics and local guarantees, not verification.

The cautionary supersets -- Objective-C, C++. Both began as C supersets. Objective-C shows one can succeed, given a backer and a clear niche; C++ shows the other outcome, where a superset accretes until it is a second, far larger language. hecc's insistence on staying small -- the two design tests, the willingness to cut features -- is a deliberate guard against the C++ trajectory.

Across all of them, hecc's slot is narrow and specific: temporal safety as an additive, erasable extension that stays ordinary C. The evolution languages ask you to switch; the spatial extensions cover a different axis; the verifiers demand proofs; the old dialects needed runtimes. hecc is the one aimed at ownership, added in place, with nothing taken away -- and, honestly, the one still on paper while most of the others already ship.

What hecc Is Trying to Do

hecc is not trying to make C safe. It is trying to make the programmer write safer C by giving them a way to explicitly communicate intent when that intent would otherwise be invisible to the compiler. Ownership transfers communicate responsibility, borrows record what a reference depends on, bounds can be checked when the compiler already knows them, and additional information can be provided when necessary. None of this needs to prevent the programmer from writing ordinary C. The difference is that hecc gives the programmer an option to say more when saying more allows the compiler to understand the program better.

There are still major problems to solve. Borrows and origins give escaping references a defined relationship to track, but the exact rules for nested and self-referential structures, for origins through function pointers, and for reconciling origins across complicated control flow still need to be worked out. External functions need contracts. Unions, realloc(), and indirect pointer manipulation all introduce additional problems. That is expected. hecc is still an experiment. The central question is whether a relatively small number of extensions can provide enough information to make static analysis significantly more useful while preserving the experience of writing C. Ownership comes first, borrows and origins follow from the problem ownership creates, and everything else should be judged by the same principle: add information where C is ambiguous, but do not add an entire new language.