Oct 9 2026 -- Eyad Amr --
GitHub
--
Discord
Disclaimer: 0% of this article was written by any tool or artificial intelligence.
Introduction
Ladies and gentlemen. Have you ever been kicked in the head by a
donkey
?
Well most of us (hopefully) have not been literally kicked in the head by any animal, but if we talk about this figuratively, you've faced the many amounts of:
For most of us; this is something we can't do much about.
But for the minority of us, and specifically the insane people like me, we get so frusturated we want to do something like building an entirely new programming language (in this case, it would be my language Zeta) that "is easier than Rust".
Obviously, if you're going to say you're going to make a language that's "easier to use than Rust while being the same performance", you'll get one of these 4 reactions:
-
"That's
extremely
ambitious."
-
"Are you out of your mind?"
-
"This wouldn't work."
-
insert 50 questions
I'm making a new language, and it's generally for a lot of reasons, but my personal favorite is to fix many of the Rust borrow checker's restrictions that can actually hurt productivity by blocking valid memory patterns and forcing the user into doing (usually more than necessary) workarounds like adding runtime checks or
unsafe
, without losing on performance guarantees
Why?
To appreciate the solution, you must first understand the problem.
You can fully skip this to the
Background
section (or even the main model itself) if you know how a Rust borrow checker works, but for the people who don't; here's a brief explanation.
The rust borrow checker is a way to write memory safe code through the sophisticated engineering of many concepts, most importantly:
-
Ownership
-
Library-discipline to create proper drops and safely abstract low-level code.
-
shared-XOR-mutable
What I exactly mean by "library-discipline to create proper drops and safely abstract low-level code." is that the compiler does not guide you towards code which would otherwise require unsafe, or runtime checks + unsafe, such as writing complex Drop implementations or creating collections.
Ownership is the action of you owning a value, such as here:
let x = String::new(); // x now owns this String.
// This `drop` call now owns `x`.
// By the end of this call, it has been freed from memory.
drop(x);
use_value(x); // compiler error: use of moved value: x
This, and combining a type of pointers called references (which is an aligned, non-null, borrow-checked pointer, but we'll exclude the borrow checked part until we actually get to the borrow checker part) singlehandedly prevents double frees, use after frees, forgetting to free (when any variable or moved function parameter reaches the end of the block, the rust compiler does this automatically) and segfaults and UB from null dereferences (you need to use an
Option
, which is a safe way to abstract it)
And this is the default model, so that means you can't implicitly copy something (collections, heap allocated memory, expensive stack allocated structs, etc) UNLESS its bits are easy to copy.
Data types that are easy and cheap to copy are pointers, immutable references (You should not be able to
Copy
&mut, even if it's cheap. Because it breaks the exclusivity guarantees it gives), any primitive and any struct that uses those things
If you don't know what a pointer is, it's basically how your computer knows where data lives in ram, so they're like addresses.
And if you're not a user of Rust, this means that you can only have 1 pointer that can be (deeply) mutated (translation to non-Rustacean: 1 pointer that can be mutated, you know your data has an address in your RAM, right? think of that address as your address. You have permitted only one to know where you live and how to enter your home, but you also gave them permission to move the furniture however you like), or infinite pointers that cannot be (deeply) mutated (translation: You have permitted infinite people to know where you live and how to enter your home, but they can only come see your home).
And if you're also not a rustacean, you may also be confused when I mention
deep
in
(deeply) mutated
. That's because if I were to do:
let mut x: Vec3f = Vec3f { x: 1.0, x: 2.0, x: 3.0 };
, then I can change
x
and get mutable "pointers" to it and change
x.x
. But if you're coming from a language like Java or C, know that if you don't add
mut
, not only can you not change
x
directly and get mutable pointers to it, but you also can'tt change
x.x
or any other field.
But pointers aren't really what we talked about, in Rust, pointers are exactly like C pointers and are therefore unsafe and dangerous and can cause many bugs. The idiomatic rust approach is using references, which is borrow checked, non-null and aligned, everything that a pointer in C and Rust has no guarantees of, so we will speak only about references unless stated otherwise.
Sounds great for a language that can be used to write such low level programs right? well the borrow checker model makes it even greater.
The borrow checker has this general rule of "shared-XOR-mutable" where you can have only one
&mut
(mutable reference) or infinite immutable references
&
.
hink of it as literally borrowing the value, you don't own it so you must return it back at some point, but you currently can use it, and they will not outlive their value, and when it dies, you essentially "give back" the access to borrow it again.
The shared-XOR-mutable is a good rule in my opinion, but Rust is so conservative it makes it exhausting to constantly satisfy it.
As a result of this restriction, rust prevents arbitrary mutations and makes mutation explicit, it is also great for
optimizations that require it so that you can only mutate a specific region of memory from one pointer, and not multiple
, notably, allowing SIMD, allowing the compiler to hold stuff in registers longer, and more.
And finally, it is used for fearless concurrency. You must either prove that you are the only holder of a mutable reference (
&mut
as we talked about, a mutable reference), it's all immutable, or you access the data which has a type of
T
, and
T
implements Send (Types that can be transferred across thread boundaries) + Sync (Types for which it is safe to share references between threads.
To understand this, we'll give this bit of code:
async fn foo() {
let mut a = 10;
let mut b = 20;
thread::scope(|s| {
s.spawn(|| {
a += 1;
});
s.spawn(|| {
b += 1;
});
});
}
Here, the closures (if you don't know what closures are, think of them like java lambdas) capture
a
and
b
as
&mut
, and since numbers are
Send
(P.S: It's a lot easier to prove references can be used in other threads, vs proving if they can be actually shared and mutated across threads) and we give exclusive references (and very important, we create a scope of threads, which joins all the threads created within that thread. If we spawned a thread directly, it may require
'static
data). But if we were to use the same value:
async fn foo() {
let mut a = 10;
thread::scope(|s| {
s.spawn(|| {
a += 1;
});
s.spawn(|| {
a += 1;
});
});
}
It will actually
not compile due to a multiple borrow error
.
The way to actually safely use
a
in parallel is to wrap it into an
Mutex<_>
like so:
async fn foo() {
let mut a = Mutex::new(10);
thread::scope(|s| {
s.spawn(|| {
*a.lock().unwrap() += 1;
});
s.spawn(|| {
*a.lock().unwrap() += 1;
});
});
}
If you wanted to not use a
scope
for whatever reason, you must at least pair it with
Arc
like so:
async fn foo() {
let a = Arc::new(Mutex::new(10));
let value = a.clone();
thread::spawn(move || {
*value.clone().lock().unwrap() += 1;
});
let value = a.clone();
thread::spawn(move || {
*value.clone().lock().unwrap() += 1;
});
}
Which atomically reference counts
a
and when the reference count hits 0, it automatically frees the memory as it's no longer used.
Well, what's to change about that?
Well, that all sounds good on paper, in reality, it's actually a lot harder to follow these rules, but very specifically
shared-XOR-mutability
First of all, it's really hard to implement observers, intrusive data structures, back-references and graphs (like doubly-linked lists), delegates, some kinds of RAII (for example, Rust cant represent a C++
struct Rollbacker { Transaction* t; ~Rollbacker() { if (t) t->rollback(); } };
), but I'd like to demonstrate why
(In Zeta, I have a third reference which is called
&alias
which allows for compiler-native, borrow checked shared mutability, and it's really ergonomic while still being completely safe from all kinds of errors, including use after frees by dangling, this "third" reference is explained later in the blog)
Anyways, here's an example:
let mut vec = Vec::with_capacity(1);
vec.push(1);
let x = vec.get(0);
vec.push(2); // cannot borrow vec as mutable because it is also borrowed as immutable. mutable borrow occurs here
println!("{}", x.unwrap());
This is actually a fair restriction, as the
Vec
can regrow and change the pointer, so
x
could be a use after free. This is a successful prevention of the borrow checker model.
But if we wanted to do something, and we have
dared
to make it more complex:
let mut vec = Vec::new();
vec.push(1);
vec.push(2);
vec.push(3);
let x = vec.get_mut(0);
let y = vec.get_mut(1); // cannot borrow vec as mutable more than once at a time. second mutable borrow occurs here
let z = vec.get_mut(2);
println!("{}", x.unwrap());
println!("{}", y.unwrap());
println!("{}", z.unwrap());
This is actually.. completely safe? Why would it reject it?
0
is clearly not
1
.. right?
Well yeah, but rust doesn't care about that.
It cares about the literal: "you can have only one
&mut
or infinite immutable references
&
", and this completely memory safe pattern is a clear violation.
Well that's fine, we have
let [x, y, z] = vec.get_disjoint_mut([0, 1, 2]).unwrap();
and we can write our code as so:
let mut vec = Vec::new();
vec.push(1);
vec.push(2);
let [x, y, z] = vec.get_disjoint_mut([0, 1, 2]).unwrap();
println!("{}", x);
println!("{}", y);
println!("{}", z);
Yes, you'd be correct. Except that
get_disjoint_mut
does an O(n^2) check, specifically, this is the rust internal check for it:
// The exact body of the function that Rust uses.
for (i, idx) in indices.iter().enumerate() {
if !idx.is_in_bounds(len) {
return Err(GetDisjointMutError::IndexOutOfBounds);
}
for idx2 in &indices[..i] {
if idx.is_overlapping(idx2) {
return Err(GetDisjointMutError::OverlappingIndices);
}
}
}
Notice how for every index, there's a completely new bound check which is 3 bound checks total, and a whopping 9 checks.
And don't forget about the error handling, the monomorphizations that happens for
iter
and
enumerate
and
..i
, and after a total of 12 checks have been done.
It proceeds to execute an unsafe function as:
// SAFETY: The `get_disjoint_check_valid()` call checked that all indices
// are disjunct and in bounds.
unsafe { Ok(self.get_disjoint_unchecked_mut(indices)) }
Well, notice how there's a comment here:
#[inline]
fn get_disjoint_check_valid<I: GetDisjointMutIndex, const N: usize>(
indices: &[I; N],
len: usize,
) -> Result<(), GetDisjointMutError> {
// NB: The optimizer should inline the loops into a sequence
// of instructions without additional branching.
...
}
So LLVM has the opportunity to fully unroll the loop as it knows the number of elements, but it still needs to do
all
of the runtime checks regardless, especially if
[0, 1, 2]
were arbitrary
[i, j, k]
where constant folding will not apply (which is the most likely case scenario)
If you're a dedicated rustacean, you may think "but the CPU can actually do pipeline and execute all comparisons in parallel by the instruction-level", That's not always true, and even if it was, it's missing the poin.
A
get_disjoint_mut
is harder to implement (it needs discipline as you're forced to write some unsafe code), and the Zeta compiler itself can sometimes
remove
entire checks.
It doesn't just know the guarantees of
get_disjoint_mut
but
literally
how it works.
The CPU would have a better time with that whether or not it can do instruction-level parallelism.
Remember when I said "it needs discipline as you're forced to write some unsafe code" for get_disjoint_mut? In my book, that's a really big problem because making it easier to write APIs, even slightly, can save a lot of time.
This is one of Rust's core limitations, the compiler may reject your code if you were to implement, say, your
Vec
.. or your
HashMap
, or your
SIMD
, or your allocators, or your collections in general, or your drop calls to free memory, or safely handling
uninitialized
memory (not null memory, there's a difference).
You usually neeed to use
unsafe
to avoid the limitations of the borrow checker and implement those drop implementations and collections. Or if you want to avoid the borrow checker
safely
, you need more runtime checks.
I think I made some really good points, it's harder to create many kinds of collections, Drop implementations, you can imagine it's already harder than normal to create production-grade software that has a lot of state.
Let's take a look of a data structure such as a
doubly linked list
that is otherwise simple in
literally
every other language: (This is not a full linked list for simplicity)
use std::{cell::RefCell, rc::Rc};
struct ListNode<T> {
item: T,
next: Option<Rc<RefCell<ListNode<T>>>>,
prev: Option<Rc<RefCell<ListNode<T>>>>,
}
impl<T> ListNode<T> {
fn new(item: T) -> Self {
Self {
item,
next: None,
prev: None,
}
}
}
#[derive(Default)]
pub struct DoublyLinkedList<T> {
head: Option<Rc<RefCell<ListNode<T>>>>;,
tail: Option<Rc<RefCell<ListNode<T>>>>;,
size: usize,
}
impl<T> DoublyLinkedList<T> {
pub fn new() -> Self {
Self {
head: None,
tail: None,
size: 0,
}
}
pub fn push_back(&mut self, item: T) {
let node = Rc::new(RefCell::new(ListNode::new(item)));
if let Some(prev_tail) = self.tail.take() {
prev_tail.borrow_mut().next = Some(Rc::clone(&node));
node.borrow_mut().prev = Some(prev_tail);
self.tail = Some(node);
self.size += 1;
} else {
self.head = Some(Rc::clone(&node));
self.tail = Some(node);
self.size = 1;
}
}
pub fn pop_back(&mut self) -> Option<T> {
self.tail.take().map(|prev_tail| {
self.size -= 1;
match prev_tail.borrow_mut().prev.take() {
Some(node) => {
node.borrow_mut().next = None;
self.tail = Some(node);
}
None => {
self.head.take();
}
}
Rc::try_unwrap(prev_tail).ok().unwrap().into_inner().item
})
}
}
impl<T> Drop for DoublyLinkedList<T> {
fn drop(&mut self) {
while let Some(node) = self.head.take() {
let _ = node.borrow_mut().prev.take();
self.head = node.borrow_mut().next.take();
}
self.tail.take();
}
}
Eugh.. you can use stuff like
type Link<T> = Option<Rc<RefCell<ListNode<T>>>>;
to replace
prev: Option<Rc<RefCell<ListNode<T>>>>,
with
prev: Link<T>
, but that's hiding the core problem.
This is a fully safe LinkedList, but it depends on several runtime checks such as:
-
Rc
for a LOT of reference counting
-
RefCell
which basically does borrow checking at runtime
-
awkward error handling such as
Rc::try_unwrap(prev_tail).ok().unwrap().into_inner().item
And we have only implemented 2 functions.
Later, I'll show the exact same implementation in Zeta, and how it is simpler.
Still not convinced about the points I've made? what about this
horrendous
lifetime syntax from just a few lines of the zeta compiler that is written in Rust?
impl<'f, 'a, 'bump> FunctionLowerer<'f, 'a, 'bump>
where
'bump: 'a,
{
'a is the general lifetime for most stuff, 'bump is the lifetime of the bump allocators and any bump allocated data, and 'f is the lifetime of the function being lowered, and how structs and usually functions have to be infested with them
pub struct FunctionLowerer<'f, 'a, 'bump> {
pub(super) current_block_data: CurrentBlockData<'f>,
pub(super) var_map: HashMap<StrId, Value>,
pub(super) phantom_data: PhantomData<&'bump ()>,
pub(super) loop_stack: Vec<LoopCtx<'a, 'bump>>,
pub(super) funcs: &'a HashMap<StrId, Function>,
pub(super) struct_field_offsets: &'a HashMap<StrId, HashMap<StrId, usize>>,
pub(super) struct_method_slots: &'a HashMap<StrId, HashMap<StrId, usize>>,
pub(super) struct_mangled_map: &'a HashMap<StrId, HashMap<StrId, StrId>>,
pub(super) struct_vtable_slots: &'a HashMap<StrId, Vec<StrId>>,
pub(super) interface_id_map: &'a HashMap<StrId, usize>,
pub(super) interface_method_slots: &'a HashMap<StrId, HashMap<StrId, usize>>,
pub(super) structs: &'a HashMap<StrId, HirStruct<'a, 'bump>>,
pub(super) enum_variant_tags: &'a HashMap<StrId, HashMap<StrId, usize>>,
pub(super) enums: &'a HashMap<StrId, HirEnum<'a, 'bump>>,
pub(super) context: Arc<StringPool>,
pub(super) extern_c_names: &'a HashSet<StrId>,
pub(super) dep_graph: &'a RefCell<DepGraph>,
pub(super) module_idx: usize,
pub(super) return_type: Option<HirType<'a, 'bump>>,
pub(super) global_funcs: &'a HashMap<StrId, Function>,
pub(super) scope_stack: Vec<DropScope<'a, 'bump>>,
pub(super) drop_state: DropMoveState<'a, 'bump>,
pub(super) glue_registry: &'a DropGlueRegistry,
pub(super) allocator_kind: &'a HashMap<StrId, AllocatorKind>,
pub(super) interface_methods: &'a HashMap<StrId, Vec<(StrId, Vec<SsaType>, SsaType)>>,
pub(super) bump: &'bump GrowableBump<'bump>,
pub(super) module_import_aliases: &'a HashMap<usize, HashMap<StrId, usize>>,
pub(super) module_named_imports: &'a HashMap<usize, HashMap<StrId, usize>>,
pub(super) constants: &'a HashMap<StrId, HirExpr<'a, 'bump>>,
pub(super) promoted_to_stack: HashSet<StrId>,
pub(super) narrowed_fields: HashMap<(StrId, Vec<StrId>), Value>,
pub(super) nullable_owned_locals: HashMap<StrId, HirType<'a, 'bump>>,
pub(super) array_flags: HashMap<StrId, (Value, usize)>,
}
And don't
ever
get me started on more complex stuff such as higher-ranked trait bounds.
I can't practically show you the
exact
equivalent of what it'd look like in Zeta, as there is no
real
software built in Zeta. There is a doubly linked list that I will be showcasing in the second half of the blog though! which uses references with
inferred
'lifetimes'.
let mut function = self
.module
.functions
.remove(&hir_fn.name)
.expect("function signature should already be registered");
let mut fl = FunctionLowerer::new(
&mut function,
In my code, this was the best way to get a
&mut
to the function I needed to lower, I needed to get the function by REMOVING it to retain ownership, and then putting it back after lowering the function.
Why couldn't I just take it? well that's because I already
have
the
self.module.functions
immutably borrowed for the FunctionLowerer to use, even though it has literally
nothing
to do with the data itself.
It's actually safe to have an immutable borrow to a struct, while having a mutable borrow to the contents of that struct, like so:
let v: Vec3f = Default::default();
let x = &mut v.x;
let y = &v;
This is basically the
example I showed
at the start of the blog.
And this same example works perfectly fine in the Zeta borrow checker. It's because of Zeta's disjointness proofs that directly make it impossible to get
&v.x
at the same time as
&mut v.x
, which I'll get to in just a second.
Background
We're done with how the borrow checker works and why it's so restrictive and conservative, now how have I even came up with this?
Well I'm actually really fresh into the memory safety models compared to many other PhDs and other researchers, but I have spent 2 consecutive years designing the perfect memory safety, from a borrow checker, and after a while, I ended up going a full circle back to a borrow checker, just not the Rust-style borrow checker.
More than a year ago, I have actually dedicated to implementing something called "compile-time refcounting", and even before that, it was "region based memory management". Both have failed in their implementations, but this borrow checker has been
fully
implemented and
fully
tested at the time of writing this, except for the fact that it doesn't support async (disclaimer: at the time of writing this)
Then on 8/28/2025, I have seen a mysterious link from a brilliant compiler engineer I've looked up to because I thought he was the most innovative of that time. It was a link for group borrowing.
https://verdagon.dev/blog/group-borrowing
I won't go too much into detail for group borrowing, so you should check it out, they are trying to achieve the same results as me in a
vastly
different way. but it seemed so
cool
, I wanted to implement it, but then I quickly asked myself, "why NOT exclusivity"?
I thought to myself, I had all the time in the world, why not try checking
borrow checkers
again? and I heavily hyperfocused on the topic
Then I started asking my self questions that today seems like it was just hiding in plain sight.
Why couldn't I just do &mut list[i] and &mut list[j] at the same time if i != j
?
I got a surge of dopamine thinking about the idea and I rushed to research it, and many months into realizing, could I make a new borrow checker entirely with no flaws?
From then, the journey began, and in 7/14/2026, almost a full year, not only had I released the first release of my language, it was also the first release to feature my
borrow checker
The fun part
We're going to start off with the base, first, it is an "shared-XOR-mutable" borrow checker, and it has 3 reference kinds. &mut, &alias and and immutable references
(fun fact: I was not programming at that time, and thought about &alias randomly while chit chatting about my language cities away from my PC and supposed to enjoy life on the beach. The more you know!)
Keep in mind the 3rd strange reference called
&alias
and the "shared-XOR-mutable" because
&alias
is about to change the meaning of "shared-XOR-mutable" as soon as we run into problems that require self-referential types, observers, etc.
'Did you just say "shared-XOR-mutable"? Well, how's that any fun? Didn't you just say that's the entire problem?'
Well, yes, but it's not the
core
problem.
The core problem is that Rust is not allowing us to exploit disjointness from a compiler native standpoint other than basic proofs like
foo.x != foo.y
which is nowhere near this borrow checker's limits.
It is safe to say that they are not giving us any breathing room, and this borrow checker will be really hard to just "contribute it" to Rust, nor can we make a compiler that transpiles to Rust (or Rust MIR)
As this is not really a superset even though they have the same "base". it's completely different in almost every aspect. I can
map
many concepts of Zeta to rust so cleanly that it could literally allow interop, but that requires sacrificing some of my other goals for my language, and that will be a later evaluation.
The rules of the borrow checker
Get prepared, because there's going to be a lot of bending and they all have reasons.
"You can't mutably borrow the same struct twice!"
Yes, except you actually can: (From this point, we will write in my language's syntax, not rust)
let mut list: ArrayList<i64> = ArrayList.new();
list.push(1);
list.push(2);
list.push(3);
x := list.get_mut(0); // Okay. This is fine!
y := list.get_mut(1); // THIS IS NOT FI- wait.. no compiler errors?
z := list.get_mut(2); // Wait.. something isn't right.. still nothing? is the compiler broken?
use_value(x); // OH NO!
use_value(y); // THE COMPILER ISN'T WORKING!
use_value(z); // CALL THE AMBULANCE!
Relax. This is valid Zeta code. and the reason is because of disjointness, and it's the subsystem that's secretly powering more than half of all Zeta's expressiveness power.
"You can't mutably borrow this because it is already immutably borrowed!"
Yes that's true, unless you mutably borrow something in the contents itself. The item stored in the collection can be mutated. The collection itself cannot, so the collection can be immutably borrowed while something in the collection itself can be mutated.
// Assume a pre-existing list and pre-existing values and pre-existing Vec3f
let x: &mut Vec3f = list.get_mut(n);
*x.x = 2;
// Rust: cannot borrow vec as immutable because it is also borrowed as mutable. immutable borrow occurs here
// Zeta: Oh yeah that's fine!
let len = list.len();
*x.y = 4;
"Lifetime doesn't live long enough"
Yes, and I don't actually have a rebuttal for this.. that's why I completely reinvented them.
We'll first talk about disjointness.
Disjointness
Disjointness is when the compiler can prove under any circumstances that the address of X does not equal the address of Y.
Back to our example:
x := list.get_mut(0);
y := list.get_mut(1);
Here, the compiler has already looked at how
get_mut
works and cached it, realizes the memory address depends on the argument we gave it, it would be capable of reasoning about pointer arithmetic directly, even if it doesn't know the result of the arithmetic, it knows the formula.
(In the case of ArrayList, it does something much better than using pointers directly)
So if we were to prove
i != j
then it would work. There are several ways to do that:
-
Constants
-
Suggestive loop indexing
-
Runtime guards
-
Arithmetic
-
Unsafe assertions (unimplemented in the compiler)
Constants:
x := list.get_mut(0);
y := list.get_mut(1); // Pass
Suggestive loop indexing:
for i in 0..n {
for j in (i + 1)..n {
x := list.get_mut(i);
y := list.get_mut(j); // Pass
}
}
Runtime guards:
i := 1;
j := get_index_from_ffi();
if (i != j) {
x := list.get_mut(i);
y := list.get_mut(j); // Pass
}
Arithmetic:
i := get_index_from_ffi();
x := list.get_mut(i);
y := list.get_mut(i + n); // pass if `n` is proven to be non-zero, if not, you need to do some sort of arithmetic, guard, etc
and lastly, unsafe assertions:
i := 1;
j := get_index_from_ffi();
unsafe { $compiler_assert(i != j); } // $function means a compiler intrinsic, in this case, we have just told the borrow checker and optimizer to assume `i != j`. This can backfire if they are equal, but that's why it's `unsafe`
x := list.get_mut(i);
y := list.get_mut(j);
If you do somethingn incorrectly, such as say, this:
let mut list: ArrayList<i64> = ArrayList.new();
list.push(1);
list.push(2);
list.push(3);
let x: &mut i64 = list.get_mut(0);
let y: &mut i64 = list.get_mut(0);
let z: &mut i64 = list.get_mut(2);
*x = 0;
*y = 0;
It gives a real compile time error:
error: cannot borrow as mutable: already mutably borrowed elsewhere (via list.data)
1685 | list.push(2);
1686 | list.push(3);
1687 | let x: &mut i64 = list.get_mut(0);
1688 | let y: &mut i64 = list.get_mut(0);
^~
1689 | let z: &mut i64 = list.get_mut(2);
1690 | *x = 0;
(Don't mind the line numbers, that's just good testing)
Interestly, if you were to delete
*x = 0;
it would compile, because NLL just erases it.
if
*y = 0;
was deleted, it did not compile.
if we changed the order of the derefs, it did not compile.
If we removed it completely, it compiled.
So NLL is fully implemented for the language, which is great!
Combined with "multi place" references (only official for functions for now)
func x(vec3f: Vec3f.{&mut x, &mut y, &z})
Can say that it mutates x and y, but only reads from
z
.
This is a core dependency of many of our safety features like
&alias
and
uninit
which would otherwise be able to unsafely bypass and create memory bugs without using
unsafe
! (We will explain this later)
It also has a really special ability, and it's the ability to directly reason about INDEXES. Meaning the borrow checker can directly benefit from:
func push_within_capacity(this.{&mut this.data[this.data.len], &mut this.data.len}) {
if (this.len == this.cap) {
debug.panic(...); // At the time of writing this it's called `debug.debug_panic`. If you are reading this, there is a chance it has changed to `panic`
}
unsafe {
// write_uninit here means that there will be no drop. `write` would do a check, assume it's contiguous and execute the drop for what it believes is already an existng value.
this.data.write_uninit(this.data.len, value); // this.data is an owned slice. it is REALLY important to the safety and ergonomics of Zeta
this.data.len += 1;
}
}
Note: this.data is an owned slice. We will get to that later in this document.
Which allows us to do something like this:
list := ArrayList::new();
list.push(1);
let x = list.get(0);
list.push_within_capacity(2);
use_value(x); // Pass! As it's proven that `push_within_capacity` only writes `this.data[this.data.len]`. there's a guarantee it does not write a new pointer to `this.data` itself like `list.push)
This subsystem can singlehandedly make or break the invention. in group borrowing's case,
group borrowing
the mental model is trying to answer what group does a reference belong to while simultaneously answering if that group still valid. In Zeta, it's more like answering the exact memory regions you are allowed to access rather than restricting exclusive mutations to memory regions that are owned by the same big memory region, but are not equal.
Owned pointers
Owned pointers and slices are basically how they sound. They are owned, which means they are subject to move semantics. But it's a pointer, so the pointer is in charge of being moved and dropped, and it's usually owned by the resource that created it.
For example, you'd create an owned pointer like this:
let node: ^Node<T> = this.alloc.new<Node<T>>(Node<T> {
value: value,
prev: null,
next: null, // Don't worry. The language has null safety, it's just a primitive instead of rust's Option
});
The main benefit is that it is automatically dropped while being always as big as a pointer, without creating entirely new library types.
Back to this example:
let node: ^Node<T> = this.alloc.new<Node<T>>(Node<T> {
value: value,
prev: null,
next: null, // Don't worry. The language has null safety, it's just a primitive instead of rust's Option
});
This would be the equivalent of this rust code:
let node: Box<Node<T>> = Box::new(Node<T> {
value: value,
prev: None,
next: None,
});
Except, we're using an allocator directly, not a
Box
that hides it away. This means that it's easier to use custom allocators:
let node: ^Node<T> = this.bump.new<Node<T>>(Node<T> {
value: value,
prev: null,
next: null,
});
This would be the equivalent of this rust code:
let node: Box<Node<T>, &Arena> = Box::new_in(Node<T> {
value: value,
prev: None,
next: None,
}, &self.arena);
Importantly, the compiler is very smart,
so
smart that it can infer that:
^Node<T>
can mean
^this.bump Node<T>
or
^this.alloc Node<T>
.
this.bump
after ^ would be the equivalent of Zeta lifetimes. and they have several qualities over rust-like lifetimes, they are called provenance (meaning: origin of place). But that is to be explained later in the document.
(the difference between a safe and unsafe pointer is that a safe pointer is non-null and aligned but not borrow checked, and an unsafe pointer is a C pointer, they are both unsafe to deref but they have different guarantees.).
And then we know what to do when the owned pointer is dropped.
In the case of a memory allocated
Player
then the player is dropped THEN freed from memory, as for the mutex, it is not dropped nor is it freed. Only the lock is dropped.
Note: Owned pointers and owned slices are zero-cost other than the fact that they get implicitly dropped, they do not affect memory usage or performance any differently than a regular pointer, and they can get enum niche optimizations and "null" optimizations
Owned slices
Let's take a look at an owned slice:
struct ArrayList<T, A: Allocator = Mallocator> {
data: ^this.alloc []T, // uninferred `alloc`
alloc: A,
}
impl<T, A: Allocator = Mallocator> ArrayList<T, A> {
func new(): ArrayList<T, Mallocator> {
let alloc: Mallocator = Mallocator {};
return ArrayList<T, Mallocator> {
data: alloc.calloc<T>(16),
alloc,
};
}
...
}
calloc
itself is safe because the pointer it drops at the end of
ArrayList
with no Drop implementation required to specifically free the pointer, the compiler is really helpful when it comes to guiding you and helping you and knowing what you want.
But owned slices are very tricky actually, because it's almost impossible to make them genuinely safe when using them for a contiguous list other than making a safe abstraction.
Mainly because it's very hard to prove whether something exists or not, you have to fill that invariant. Which is why
write_uninit
would be needed here:
func push_within_capacity(this.{&mut this.data[this.len], &mut this.len}) {
if (this.len == this.cap) {
debug.panic(...);
}
unsafe {
this.data.write_uninit(this.data.len, value);
this.data.len += 1;
}
}
Here, we specifically do not want to call a drop, if we used
this.data[this.data.len] = value;
the compiler will refuse it if you didn't prove under any circumstances that there may be a real value there.
mutation of
this.data.len
is also unsafe, because well.. that's a logic bug, and it indirectly causes memory bugs, but it's still a logic bug at the end of the day. and it's very hard to prove logic bugs are actually bugs because the compiler only knows what's definitely safe and what's not, so that's unsafe.
It also comes with
three implicit fields
, not two like a borrowed slice (ptr,
len
that can be accessed but not mutated). an owned slice comes with
ptr
,
len
(unsafe to mutate) and
cap
(the amount of allocated memory in there, not mutable)
The difference is although, it's a lot
easier
to use an owned slice to do what you want than a pointer, because you don't need to handwrite pointer arithmetic, deal with unsafe allocators or write your own drop implementations.
Provenance
Provenance is Zeta's solution to all the horrible things we talked about Rust's lifetimes.
It does not just say how long a reference/owned pointer is allowed to live, it also says how a certain reference/owned pointer is owned.
For example, if you look at the error message from a while back:
cannot borrow as mutable: already mutably borrowed elsewhere (via list.data)
list.data is the provenance that
x := list.get_mut(0);
has to go through, so
x
has the same maximum lifetime of as
list.data
, not
list
, in other words,
x
can be uninferred to
&list.data mut i64
.
This has a few benefits:
Can be inferred a lot easier
This allows the provenance of a reference or an owned pointer to be determined by how it is used, not by its general lifetime, that way, you know if
list.data
dies before the reference itself dies instead of some arbitrary
'a
that may or may not be inferred by the mercy of the Rust compiler.
If we go to peek the LinkedList's declaration, we can find this piece of code:
struct Node<T> {
value: T,
prev: ?&alias Node<T>, // Am I tripping? or does it not have any lifetimes?
next: ?^Node<T>,
}
struct LinkedList<T, A: Allocator = Mallocator> {
head: ?^Node<T>,
tail: ?&alias Node<T>,
len: usize,
alloc: A,
}
It's easy to infer the lifetime of the reference from this code:
impl<T, A: Allocator = Mallocator> LinkedList<T, A> {
func new(): LinkedList<T, Mallocator> {
let alloc: Mallocator = Mallocator {};
return LinkedList<T, Mallocator> {
head: null,
tail: null,
len: 0,
alloc,
};
}
...
}
And then:
this.head = node;
this.tail = &alias *this.head;
No matter what allocator you give, it's all the same anyways, the allocator owns
head
and
next
, and
tail
/
prev
specifically borrows that data. This is what I mean when I say provenance can be inferred based on the usage of the references or how you actually create those references/owned pointers
Better diagnostics
I mean come on.. which diagnostic will you want to see?
"Lifetime does not live long enough" with some lifetime hints?
or:
"this.alloc does not live as long as
this
" with locations to when it's dropped, the field and the associated function that deals with
this.alloc
improperly.
Safe to say unless you are masochist and like pain and suffering, you want the latter.
Sometimes, Rust can give good diagnostics in the case of:
fn main() {
let r;
{
let x = 42;
r = &x;
}
println!("{r}");
}
giving you:
error[E0597]:
x
does not live long enough
It's usually a lot more complex though, as real software are a lot less trivial to reason about.
Owned pointers and owned slices depend on it
Owned pointers and slices depend on the provenance (but specifically an owner provenance) to know if that type can be used to allocate the owned pointer and then drop the owned pointer automatically.
So even if we were to add real lifetimes like Rust for
any
reason, this would still exist. it's just that provenance (till now) does its lifetime work
so
good that it could theoretically rust-like lifetimes could be deemed useless for most cases (This is a
very
bold claim for something theoretical, and would require real software to be built into Zeta to stress test just how true this fact is)
&alias
&alias is a little different than the "base" we talked about.
remember when I said '
&alias
is about to change the meaning of "shared-XOR-mutable" really quick', well,
&alias
allows you to hold more than 1 mutable alias to the same data.
The problem with all that we discussed with shared-XOR-mutable memory paradigm is that we still have no real way to model many kinds of graphs.
If we go back to the rust DoublyLinkedList example we can find this:
pub fn push_back(&mut self, item: T) {
let node = Rc::new(RefCell::new(ListNode::new(item)));
if let Some(prev_tail) = self.tail.take() {
prev_tail.borrow_mut().next = Some(Rc::clone(&node));
node.borrow_mut().prev = Some(prev_tail);
self.tail = Some(node);
self.size += 1;
} else {
self.head = Some(Rc::clone(&node));
self.tail = Some(node);
self.size = 1;
}
}
Specifically, we can find that there's a chance that
self.head == self.tail
, meaning
&mut
is actually incompatible, because the stable linked list requires that the nodes
might
not
be exclusive.
How do we fix this without reimplementing rust interior mutability and refcounting that interior mutability anyway OR using pointers of any kind? That would be
&alias
. Which is not exclusive but you can mutate through it.
Which means:
x := &alias arr[0];
y := &alias arr[0];
is actually
valid
and will compile. It obviously comes with the downsides that it won't get the
noalias
guarantees and you likely won't be able to share it across threads at all unless if
T
in
&alias T
is Send + Sync.
"Well I have a really good example of why this is actually BAD and not memory safe"
I know, I know. You're about to argue how this is basically the literal definition of an
UnsafeCell
, but it's not.
First of all,
&alias
is incompatible with an exclusive reference (due to exclusivity) and with an immutable reference (to retain immutability guarantees safely), by opting into
&alias
you specifically agree that anything can mutate your stored value in any way. and that can make logic bugs easier to create, but its benefits vastly outweigh that slightly bigger possibility, in an addition, it's still completely memory safe, you won't be able to dangle it, move the value it references, or make it do some sort of UB
"Actually, you can dangle it!"
How?
A dangling possibility
This skeptical user has raised a really smart question. How can this:
let x = vec.get_alias(1);
vec.push(5);
// Can you use x here because it has no guarantee of exclusivity?
Be safe?
Well simple. vec.push would require a &mut, and &alias is incompatible with &mut and &, and &mut is incompatible with any reference to the same memory. So you can only have one reference
That alone actually still allows UB, because a not-so-friendly programmer can make their own collection with
push
taking in an &alias.. so surely that's unsafe right?
Yes! but not really because of an important subsystem that prevents this.. the core of this entire borrow checker
How disjointness singlehandedly fixes this indirectly
It actually
doesn't
let you. And it gives you this:
error: the reference's lifetime has ended because it was invalidated (via value)
2012 | // The fact that `value` is &alias rather than &mut must NOT allow
2013 | // the dangling reference to be used.
2014 | //
2015 | let stale: i64 = *value;
^~~~~~
2016 |
2017 | return 0;
Let's demonstrate:
func test_array_list_alias_invalidation_after_push(): i64 {
io.stdout().writeln("=== Running: test_array_list_alias_invalidation_after_push ===");
let mut list: ArrayList<i64> =
ArrayList.with_capacity(2);
list.push(10);
list.push(20);
// get_mut() returns &alias i64.
// This reference is allowed to coexist with other shared accesses,
// but it must not survive an operation that can invalidate it.
let value: &alias i64 = list.get_mut(0);
io.stdout().writeln("[PASS] obtained &alias reference");
// This push exhausts the capacity and causes ArrayList to grow.
// The old allocation is therefore invalidated.
list.push(30);
io.stdout().writeln("[PASS] push() completed");
// EXPECTED BORROW-CHECKER ERROR:
// `value` refers into the old ArrayList allocation, which was
// invalidated by push().
//
// The fact that `value` is &alias rather than &mut must NOT allow
// the dangling reference to be used.
//
let stale: i64 = *value;
return 0;
}
To explain this thoroughly, we should go back to multi-place references.
Remember when we said we had "multi-place references" like
this.{&mut this.data[this.len], &mut this.data.len}
? Well, the compiler can actually infer that into
&mut this
. If we were to do:
&alias this
then the borrow checker would expand this to "
this.{&alias this.data, &alias this.data.len}
"
and guess what, if you invalidate
&alias this.data
, the compiler catches that you can do
&alias this.data[i]
, that would be invalidated.
That actually fixes the problem. And specifically it's because the borrow checker
will
infer it.
And you can't actually lie because the type checker and borrow checker work together to see if what you did is correct, so that syntax can work just for the sake of explicitness, and any potential uses that an API or a closure/lambda (where we can't see its body, so we have to rely on the signature) will need
So no,
&alias
is safe. And if we remove this test, the same
LinkedList
that fully depends on this, works completely fine!
the LinkedList and implementation details
If you've reached this far, you seem interested in how this all picks up.
Here's an example (and as of the writing of this document, it is currently the main LinkedList in the standard library)
package zeta::utils::linked_list;
import zeta::alloc::allocator.Allocator;
import zeta::utils::mallocator.Mallocator;
import zeta::io;
struct Node<T> {
value: T,
prev: ?&alias Node<T>,
next: ?^Node<T>,
}
struct LinkedList<T, A: Allocator = Mallocator> {
head: ?^Node<T>,
tail: ?&alias Node<T>,
len: usize,
alloc: A,
}
impl<T, A: Allocator = Mallocator> LinkedList<T, A> {
func new(): LinkedList<T, Mallocator> {
let alloc: Mallocator = Mallocator {};
return LinkedList<T, Mallocator> {
head: null,
tail: null,
len: 0,
alloc,
};
}
func push_back(&mut this, value: T) {
let node: ^Node<T> = this.alloc.new<Node<T>>(Node<T> {
value: value,
prev: null,
next: null,
});
this.len += 1;
if (this.tail == null) {
this.head = node;
this.tail = &alias *this.head;
return;
}
let tail: ?&alias Node<T> = this.tail;
node.prev = tail;
tail.next = node;
this.tail = tail.next;
}
func remove_back(&mut this): ?T {
let tail: ?&alias Node<T> = this.tail;
if (tail == null) {
return null;
}
let prev: ?&alias Node<T> = tail.prev;
this.len -= 1;
if (prev != null) {
let node: ?^Node<T> = $replace(prev.next, null);
this.tail = prev;
if (node == null) {
$unreachable();
}
return node.value;
} else {
let node: ?^Node<T> = $replace(this.head, null);
this.tail = null;
if (node == null) {
$unreachable();
}
return node.value;
}
}
}
Let's compare the 2 languages in how they implement the same thing:
use std::{cell::RefCell, rc::Rc};
struct ListNode<T> {
item: T,
next: Option<Rc<RefCell<ListNode<T>>>>,
prev: Option<Rc<RefCell<ListNode<T>>>>,
}
impl<T> ListNode<T> {
fn new(item: T) -> Self {
Self {
item,
next: None,
prev: None,
}
}
}
#[derive(Default)]
pub struct DoublyLinkedList<T> {
head: Option<Rc<RefCell<ListNode<T>>>>,
tail: Option<Rc<RefCell<ListNode<T>>>>,
size: usize,
}
impl<T> DoublyLinkedList<T> {
pub fn new() -> Self {
Self {
head: None,
tail: None,
size: 0,
}
}
pub fn push_back(&mut self, item: T) {
let node = Rc::new(RefCell::new(ListNode::new(item)));
if let Some(prev_tail) = self.tail.take() {
prev_tail.borrow_mut().next = Some(Rc::clone(&node));
node.borrow_mut().prev = Some(prev_tail);
self.tail = Some(node);
self.size += 1;
} else {
self.head = Some(Rc::clone(&node));
self.tail = Some(node);
self.size = 1;
}
}
pub fn pop_back(&mut self) -> Option<T> {
self.tail.take().map(|prev_tail| {
self.size -= 1;
match prev_tail.borrow_mut().prev.take() {
Some(node) => {
node.borrow_mut().next = None;
self.tail = Some(node);
}
None => {
self.head.take();
}
}
Rc::try_unwrap(prev_tail).ok().unwrap().into_inner().item
})
}
}
impl<T> Drop for DoublyLinkedList<T> {
fn drop(&mut self) {
while let Some(node) = self.head.take() {
let _ = node.borrow_mut().prev.take();
self.head = node.borrow_mut().next.take();
}
self.tail.take();
}
}
As you can see, they are both fundamentally different. To a regular human, Rust looks more concise and more per-line dense. That's honestly the fault of Zeta, as it's still an experimental language and so has less polish on how good the diagnostics and generic ergonomics (not the borrow checker ergonomics, the borrow checker is
fully
implemented and I'll go in depth on that.), this will be improved later with better expressions, 1 line null handling, and more.
But if we focus on what the language does exactly, there's a massive difference and it's not even close.
First, Zeta's LinkedList has no Drop implementation. That is actually intentional, as Zeta knows how to drop self-referential types, as per this test:
struct DroppableItem {
tag: str,
}
impl DroppableItem by Drop {
func drop(mut this): void {
let mut out: FileWriter = io.stdout();
out.writeln(" -> [DROP] ArrayList DroppableItem");
out.writeln(" tag:");
out.writeln(this.tag);
}
}
func test_linked_list_dropping_test(): i64 {
io.stdout().writeln("=== Running: test_linked_list_dropping_test ===");
let mut list: LinkedList<DroppableItem> = LinkedList.new();
list.push_back(DroppableItem { tag: "clear_one" });
list.push_back(DroppableItem { tag: "clear_two" });
list.push_back(DroppableItem { tag: "clear_three" });
io.stdout().writeln("Should see 3 drops after this println");
return 0;
} // It's ran in `main` afterwards.
=== Running: test_linked_list_dropping_test ===
Should see 3 drops after this println
-> [DROP] ArrayList DroppableItem
tag:
clear_one
-> [DROP] ArrayList DroppableItem
tag:
clear_two
-> [DROP] ArrayList DroppableItem
tag:
clear_three
So the Drop implementation works correctly, if you notice the 2 remove_back/pop_back implementations though:
pub fn pop_back(&mut self) -> Option<T> {
self.tail.take().map(|prev_tail| {
self.size -= 1;
match prev_tail.borrow_mut().prev.take() {
Some(node) => {
node.borrow_mut().next = None;
self.tail = Some(node);
}
None => {
self.head.take();
}
}
Rc::try_unwrap(prev_tail).ok().unwrap().into_inner().item
})
}
func remove_back(&mut this): ?T {
let tail: ?&alias Node<T> = this.tail;
if (tail == null) {
return null;
}
let prev: ?&alias Node<T> = tail.prev;
this.len -= 1;
if (prev != null) {
let node: ?^Node<T> = $replace(prev.next, null);
this.tail = prev;
if (node == null) {
$unreachable(); // This panics btw.
}
return node.value;
} else {
let node: ?^Node<T> = $replace(this.head, null);
this.tail = null;
if (node == null) {
$unreachable();
}
return node.value;
}
}
There's a clear winner. The Rust version goes through reference counting, runtime checks from RefCell (because it's validated at RUNTIME. not compile-time like
&alias
), and to be fair, Zeta here is a lot more explicit about what it does, it feels less like an easier Rust and more like a memory safe C, if that makes sense, so already on paper, if they were both compiled with LLVM, Zeta should be faster, because it can exploit
more
optimizations (due to disjointness), require less runtime checks and still get noalias on every single
&mut
. Here there's no noalias for
&alias
because.. well, it just isn't noalias. Lol
There's also a massive difference between the fields:
next: Option<Rc<RefCell<ListNode<T>>>>,
prev: Option<Rc<RefCell<ListNode<T>>>>,
prev: ?&alias Node<T>,
next: ?^Node<T>,
way
shorter. It's also obvious what the relationships are, the
head
and
next
own everything,
prev
and
tail
borrow, and life goes on!
Async borrow checking.
Async borrow checking (as of this blog) is the same as Rust. You need
Send
and
Sync
,
&alias
requires you to be
&alias
, and threads that are unscoped take references with
&static
provenance or owned values, or
Arc
.
But async borrow checking benefits from our disjointness system like here:
func test_scope_four_disjoint_slices(): i64 {
io.stdout().writeln("=== Running: test_scope_disjoint_slices ===");
let mut level: Level = Level { tiles: undefined };
let N: i64 = 64;
let done: i64 = scope(func(s: &alias Scope) {
let a: &mut []Tile = &mut level.tiles[0..<(N / 4)];
let b: &mut []Tile = &mut level.tiles[(N / 4)..<(N / 2)];
let c: &mut []Tile = &mut level.tiles[(N / 2)..<(3 * N / 4)];
let d: &mut []Tile = &mut level.tiles[(3 * N / 4)..<N];
s.spawn(func() {
fill_tiles(a, 1);
return 0;
});
s.spawn(func() {
fill_tiles(b, 2);
return 0;
});
s.spawn(func() {
fill_tiles(c, 3);
return 0;
});
s.spawn(func() {
fill_tiles(d, 4);
return 0;
});
return 0;
});
// The scope has joined all four threads.
if (level.tiles[0].kind != 1 || level.tiles[15].kind != 1) {
io.stdout().writeln("[FAIL] first quarter not filled");
return 1;
}
if (level.tiles[16].kind != 2 || level.tiles[31].kind != 2) {
io.stdout().writeln("[FAIL] second quarter not filled");
return 2;
}
if (level.tiles[32].kind != 3 || level.tiles[47].kind != 3) {
io.stdout().writeln("[FAIL] third quarter not filled");
return 3;
}
if (level.tiles[48].kind != 4 || level.tiles[63].kind != 4) {
io.stdout().writeln("[FAIL] fourth quarter not filled");
return 4;
}
io.stdout().writeln("[PASS] scope with four disjoint &mut slices");
return 0;
}
Which compiles perfectly fine, but if you try to be sneaky by changing
&mut level.tiles[(N / 2)..<(3 * N / 4)]
of
c
to
&mut level.tiles[(N / 2)..<(4 * N / 4)]
, the borrow checker does catch it. If you violate the borrow checker, you may get errors for this code:
func test_scope_disjoint_slices(): i64 {
io.stdout().writeln("=== Running: test_scope_disjoint_slices ===");
let mut level: Level = Level { tiles: undefined };
let done: i64 = scope(func(s: &alias Scope) {
let a: &mut []Tile = &mut level.tiles[0..=16];
let b: &mut []Tile = &mut level.tiles[16..=32];
s.spawn(move func() { fill_tiles(a, 1); return 0; });
s.spawn(move func() { fill_tiles(b, 2); return 0; });
return 0;
});
// The scope has joined both threads, so `level` is usable again.
if (level.tiles[0].kind != 1 || level.tiles[15].kind != 1) {
io.stdout().writeln("[FAIL] first half not filled by thread a");
return 1;
}
if (level.tiles[16].kind != 2 || level.tiles[31].kind != 2) {
io.stdout().writeln("[FAIL] second half not filled by thread b");
return 2;
}
if (level.tiles[32].kind != 0) {
io.stdout().writeln("[FAIL] untouched region was modified");
return 3;
}
io.stdout().writeln("[PASS] scope with disjoint &mut slices");
return 0;
}
Like so:
error: cannot borrow as mutable: already mutably borrowed elsewhere at line number 3333 at column 52 inside of file named ./src/input.zeta
3330 |
3331 | let done: i64 = scope(func(s: &alias Scope) {
3332 | let a: &mut []Tile = &mut level.tiles[0..=16];
3333 | let b: &mut []Tile = &mut level.tiles[16..=32];
^~~
3334 |
3335 | // `move` is required: `a` and `b` are locals of this closure, and the
error: cannot prove these two accesses don't overlap (via b) at line number 3338 at column 42 inside of file named ./src/input.zeta
3335 | ...
3336 | ...
3337 | s.spawn(move func() { fill_tiles(a, 1); return 0; });
3338 | s.spawn(move func() { fill_tiles(b, 2); return 0; });
^~
3339 |
3340 | return 0;
error Found 2 errors
and yes, if I were to do:
let a: &mut []Tile = &mut level.tiles[0..=16];
let b: &mut []Tile = &mut level.tiles[17..=32];
or
let a: &mut []Tile = &mut level.tiles[0..=16];
let b: &mut []Tile = &mut level.tiles([16 + 1)..=32];
it compiles, while
17 - 1
gives you this:
error: cannot borrow as mutable: already mutably borrowed elsewhere at line number 3333 at column 58 inside of file named ./src/input.zeta
3330 |
3331 | let done: i64 = scope(func(s: &alias Scope) {
3332 | let a: &mut []Tile = &mut level.tiles[0..=16];
3333 | let b: &mut []Tile = &mut level.tiles[(17 - 1)..=32];
^~~
3334 |
3335 | // `move` is required: `a` and `b` are locals of this closure, and the
error: cannot prove these two accesses don't overlap (via b) at line number 3338 at column 42 inside of file named ./src/input.zeta
3335 | ...
3336 | ...
3337 | s.spawn(move func() { fill_tiles(a, 1); return 0; });
3338 | s.spawn(move func() { fill_tiles(b, 2); return 0; });
^~
3339 |
3340 | return 0;
error Found 2 errors
Bonus: Uninitialized memory
Before continuing, know that uninitialized memory does not equal null memory in Zeta.
[ T ][ T ][ T ][ ? ][ ? ][ ? ][ ? ] ...
^ ^ ^
valid values
The first three locations contain actual T values.
The remaining locations are allocated memory (it does not matter whether it's on the stack or heap), but there is no T there yet. This is unlike nullables in Zeta because this is not checked. If you access an uninitialized variable, then it's UB.
This is useful for things like ArrayList; we don't want to construct 16 values just because the list has capacity for 16 values, which can add additional memory safety for no reason.
But it has its problems, and we'll first talk about the problems that can't really be solved:
Take this ArrayList declaration for example:
struct ArrayList<T> {
data: ^this.alloc []T,
len: usize,
}
and:
list.push(value);
then eventually we need to do something equivalent to:
data[len] = value;
len += 1;
But there is an important problem.
A normal assignment may assume that there's an already existing value and therefore the compiler would generate a drop call.. except there's no existing value, if you were to be able to inflict pain on someone by bonking them in the air with nothing in your hand, you would have to do a lot of impossible explanations on how that's physically possible.
So in summary, treating this as an ordinary assignment would be incorrect.
Putting
uninit
can make it more complex than manually handling the invariant ourselves, because compared to something like:
func from_str(path: str): CPathBuf {
let bytes: &[]u8 = path.as_bytes();
if (bytes.len >= 4096) {
debug.debug_panic("path exceeds max length");
}
let mut cbuf: CPathBuf = CPathBuf { buf: uninit, len: bytes.len };
for (let i: usize = 0; i < bytes.len; i += 1) {
cbuf.buf[i] = bytes[i];
}
cbuf.buf[bytes.len] = 0;
return cbuf;
}
Where
uninit
here can be trivially reasoned about, sometimes it can be really hard to do it while trying to assume an ArrayList can handle heavy conditionals, arbitrary non-deterministic mutation (you can't predict when a User will execute a command that appends some data to an ArrayList), and more risks, when we could just add 1 unsafe method (the
write_uninit
function) and manually handle that the invariant is that the owned slice (which theoretically is an unsafe
uninit
array) is contiguous, which means there can't be something like:
[ T ][ T ][ T ][ ? ][ T ][ ? ][ ? ] ...
^ ^ ^ ^
valid values
Which would break the whole purpose of an ArrayList.
But if we exclude that case, then what's the purpose of me including this section?
Well, if we go back to our example:
func from_str(path: str): CPathBuf {
let bytes: &[]u8 = path.as_bytes();
if (bytes.len >= 4096) {
debug.debug_panic("path exceeds max length");
}
let mut cbuf: CPathBuf = CPathBuf { buf: uninit, len: bytes.len };
for (let i: usize = 0; i < bytes.len; i += 1) {
cbuf.buf[i] = bytes[i];
}
cbuf.buf[bytes.len] = 0;
return cbuf;
}
You might notice that.. this actually needs
uninit
memory, so what
is
uninit? It's uninitialized memory!
in C talk, this:
equals this:
And if you were to access
x
in C, it is UB. It may straight up segfault, read nothing, and even open up security vulnerabilities if done wrong!
But you can't use
uninit
like this in Zeta, while still being really ergonomic in the safe case. And it's all thanks to the borrow checker!
How this borrow checker model helps me safely define
uninit
.
First of all, the main limitation of
uninit
is that you can't access a piece of memory of X before it's been initialized. That seems fair, but exactly how does it work?
Let's see:
let x: Vec3f = uninit; // Assuming Vec3f has 3 f32 fields.
// This does not compile
//use(x);
x.x = 5.0f;
// This still does not compile
//use(x);
use_f32(x.x); // Pass
// This also does not compile
//use_f32(x.z);
x.y = 5.0f;
x.z = 5.0f;
use_f32(x.y); // Pass
use_f32(x.z); // Pass
use(x); // Pass
It works at a field-level for structs, and a whole-value for numbers. But it can also work for arrays:
let mut arr: [3]DroppableItem = uninit;
// Compiler knows that no existing value exists.
arr[0] = DroppableItem { tag: "arr_0_original" };
// Compiler knows that no existing value exists.
arr[1] = DroppableItem { tag: "arr_1_original" };
// Compiler knows that no existing value exists.
arr[2] = DroppableItem { tag: "arr_2_original" };
// Compiler knows that an existing value DOES exist.
arr[0] = DroppableItem { tag: "arr_0_replacement" };
// Here, we proved all of 0..<3 is initialized. So thiss is valid
let s: &mut []DroppableItem = &mut arr[0..<3];
// The compiler will always assume that there are old values in a slice.
s[1] = DroppableItem { tag: "arr_1_via_slice" };
This also works for loops:
let mut cbuf: CPathBuf = CPathBuf { buf: uninit, len: bytes.len };
for (let i: usize = 0; i < bytes.len; i += 1) {
cbuf.buf[i] = bytes[i];
}
cbuf.buf[bytes.len] = 0;
Here, we don't
know
what bytes.len, all we know is that it's a maximum of 4096, and even if
bytes.len
were to be 0, we know that at least, the first element is always initialized, and it's literally just a null termination, which the OS will atomically validate when doing:
let cpath: CPathBuf = CPathBuf.from_str(path);
let fd: i32 = unsafe { libc.open(cpath.ptr(), flags, mode) };
So it's always a safe operation.
Quirks that are fixed
First, you cannot uninitialize something that's already been initiialized. That's fair, it allows this feature to stay deterministic instead of trying to solve arbitrary runtime logic which can sometimes be impossible to reason about.
There's a quirk that's harder to fix though; What if I were to initialize an uninitialized value through something seemingly opaque?
We'd need to know how opaque. There are different levels of it:
Through a regular function call
We first need to establish that you cannot give immutable references of partially-initialized or uninitialized memory to any function. But we can allow them to be mutated, and if there's a possibility where the function reads the value before mutating it, the call would be rejected.
Such as so:
let mut cell: i64 = uninit;
fill(&mut cell);
Well, how would we prove this you may ask? Well, you tell me. I've already explained it to you like three times, you seem to forget a lot!
this.{&mut this.data[this.data.len], &mut this.data.len}
multiplace syntax. For here,
fill
would be:
func fill(x: &mut i64) {
*x = 7;
}
It takes the pointer as whole, and initializes it whole. that's valid. But can we lie and just.. not initialize it anyway? we can do:
case scenario #1:
func fill(x: &mut i64) {
}
or case scenario #2:
func fill(x: &mut i64) {
if random(1, 10) == 5 {
*x = 7;
}
}
The compiler is really smart. because it has already analyzed
fill
. so case #1 fails and the borrow checker bonks you on the head for doing so.
As for
fill
, it's a little harder than that. We usually deny the usage of the variable unless all branches (the if branch, and what if the if branch didn't happen) initialized the variable. Because what if we were to do:
func fill(x: &mut DroppableItem) {
if random(1, 10) == 5 {
*x = DroppableItem { tag: "x" };
}
}
The problem here is that we can genuinely skip a Drop call here.
Well this is where the compiler puts its foot down, and it allows it. But it will implicitly allocate a small number flag that says if the variable is initialized or not. Should this be made explicit by a compiler highlight and tooling perhaps? it's better than nothing, honestly.
Through an opaque dynamic dispatch call or opaque closure/lambda/function pointer
Take something like:
let x: Vec3f = uninit;
let mutator: &dyn MyTrait = get_my_interface(7);
mutator.mutate(x);
This is more complicated, because we don't know which implementation we actually got.
We can just reject the call conservatively, but we can allow the call
if
MyTrait#mutate
says exactly what it does via the multi-place syntax on the signature itself.
This requires it so that the implementors of
MyTrait
must follow the exact orders its been given.
So if an implementor were to not initialize all of
Vec3f
and
MyTrait
allowed
&mut X
, the call to
mutate
with an uninitialized variable, and there would be a diagnostic saying that "Possible implementors such as [
EvilMyTraitImplementation
] have different guarantees opposed to [
NotAnEvilMyTraitImplementation
]. Uninitialized memory cannot be given.". This is obviously over-restrictive if we know we'd never get
EvilMyTraitImplementation
. But honestly? That's just life. And there would be no way to prove that unless we'd be willing to implement something even more complex than what we already have.
If the
MyTrait#mutate
were to be:
func mutate(&this, value: Vec3f.{&mut x});
Then every implementation of the interface can only write to
&mut x
and never write more or less. Therefore it's perfectly fine to give uninitialized memory and always expect Vec3f.x to be usable.
It would be very similar with safe function pointers (excluding JIT functions, functions that follow C ABI, functions that come from FFI, etc), closures and lambdas
Opaque FFI functions
Uninitialized memory cannot be proved to be initialized or stay uninitialized after an FFI call. Therefore you wouldn't be able to
safely
assume it would initialize.
If you were to opt into
unsafe
though. It would be different:
let x: i64 = uninit;
let y: [*]mut i64 = &mut x;
unsafe { some_ffi_function(y); }
use_i64(x);
some_ffi_function
here doesn't take a reference. it takes an unsafe pointer.
Remember when I mentioned "the difference between a safe and unsafe pointer is that a safe pointer is non-null and aligned but not borrow checked, and an unsafe pointer is a C pointer, they are both unsafe to deref but they have different guarantees."
Well,
*mut i64
would be a safe pointer, and
[*]mut i64
would be an unsafe pointer. In this case, we use the
[*]mut i64
as we have no other way to prove that it would retain the safe pointer's (let alone the reference's) guarantees of alignment and nullability.
Using
x
afterwards would be considered legal. Why wouldn't we just make it unsafe to dereference the value and banning the ability to build a reference to
x
without unsafe? instead of potentially lying? because otherwise,
uninit
is harder outside FFI. I think this would be fine though so I'd definitely like more feedback regarding this.
Static memory
Static memory is really easy to explain,
&mut
to static is banned as it's almost impossible to track, and you'd use
&alias
which has no guarantee of data staying the same always and is never
noalias
which is correct, and static memory lives as long as the program, and values do not require
Send
+
Sync
unless it crosses thread boundaries, nor does it need
unsafe
to use.
Conclusion
We've finished the full Zeta borrow checking model. This is obviously way easier said than done, but there's a real implementation that actually implements this. Which is.. you guessed it. Zeta!
Again,
this here
is the github and
this here
is the discord link for anyone who wants to chat about this model!