I believe this is barking up the wrong tree since IMHO C is just "high level assembly" for systems programming. As soon as you add runtime behavior to combat Undefined Behavior (UB) you're blowing up execution times. And static analysis can only go so far without blowing up compile times.
C is "the right tool for the right job" which is operating systems and its code which is called thousands of times per second. You cannot afford even one iota of runtime checks in that code. The developer must know what he's doing or he should get out of the kitchen.
We should discourage the usage of C in application programming and prod developers towards memory safe languages like Rust or Go.
And I'm not even sure if Rust solves this case as far as UB is concerned.
"Some Honeywell machines, for example, had nine-bit bytes."
OS 2200 has 36-bit words. It is still a supported platform.
https://en.wikipedia.org/wiki/UNIVAC_1100/2200_series
This platform was the first SMP UNIX implementation:
"Any configuration supplied by Sperry, including multiprocessor ones, can run the UNIX system."
https://www.nokia.com/bell-labs/about/dennis-m-ritchie/other...
1790620504 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/b9325731ea9b61e0/ | https://news.ycombinator.com/item?id=49882419 | 15 comments
1790673092 | Reducing undefined behavior in the C language | https://lwn.net/SubscriberLink/1095811/efcdbcf080cfa4c6/ | https://news.ycombinator.com/item?id=49890290 | 0 comments
I recently had a heated but productive discussion about undefined behavior here.
As per the linked article:
“There are currently about 100 instances of undefined behavior in the C standard, but the in-progress C2y draft has removed 45 of them.”
I wonder how they handle the specific case of uninitialized but allocated memory.
Let’s look at something which will result in undefined behavior in C99: [1]
#include<stdio.h>
#include<stdint.h>
#include<stdlib.h>
#define b(z) for(c=0;c<z;c++)
uint32_t c,e[42],f[42],g=19,h
=13,n[45],i,j,k;void m(){j=0;
b(12)f[c+c%3*h]^=e[c+1];b(g){
i=c*7%g;k=e[i++];k^=e[i%g]|~e
[(i+1)%g];j=j+c;n[c]=n[c+g]=k
>>j%32|k<<-j%32;}for(i=39;i--
;f[i+1]=f[i])e[i]=n[i]^n[i+1]
^n[i+4];b(3)e[c+h]^=f[c*h]=f[
c*h+h];*e^=1;}int main(int c,
char**v){char*q=malloc(2);if(
q==0)return 0;q[0]&=31;q[0]|=
64;q[1]=0;for(;;m()){b(3){
for(j=0;j<4;){f[c*h]^=k=(*q?
255&*q:1)<<8*j++;e[c+16]^=k;
if(!*q++){b(18)m();b(2){j=c;
b(4)printf("%02x",(e[1+j%2]
>>8*c)&255);c=j;if(c%2)m();}
puts("");return 0;}}}}}
The key part of the above brick of code is this: char *q=malloc(2);
if(q==0)return 0;
q[0]&=31;
q[0]|=64;
q[1]=0;
Here, we see that q[0] is an allocated but undefined byte. As per C99, this results in undefined behavior, however 20 years ago this was a good trick to get kinda-randomish bytes to use as a possible entropy source.Someone claimed that the above brick of code will compile in newer versions of clang such that, since the complex cryptographic pseudo random number generator code depends on uninitialized but allocated memory, the entire cryptographic operation isn’t performed.
So I tested it against multiple versions of GCC and clang; I also tested it against TCC for good measure.
In all cases, with all levels of optimization, the cryptographic routine ran. I even ran it against clang 23. In cygwin, it was a randomish but consistent byte (except for clang at a higher level of optimization, at which point the uninitialized byte had a value of 0); in Ubuntu 26, the uninitialized memory consistently had a value of 0 (in tcc/gcc/clang).
I am hoping the up and coming C2y spec has very clearly defined behavior when using unintialized memory (ideally where it will work but the bytes can have any values).
Naturally, I have updated my code to no longer use uninitialized memory as a source of entropy. 20 years ago, MacOS didn’t support clock_gettime() with nanosecond resolution, so that wasn’t a portable way to get pseudo-random bits; these days clock_gettime() is universal across modern development environments, and it provides pretty good entropy (along with using /dev/urandom in *NIX, which isn’t in POSIX but is widely supported, as well as CryptGenRandom() in the legacy Win32 port). [2]
[1] Said person said the appendices to C99 aren’t authoritative, but if something is in the spec, including in the appendices, it’s authoritative.
[2] I don’t blindly trust /dev/urandom to always make really hard to guess pseudo-random bits, because my code is open source, and, as such, doesn’t just compile in Linux. It often times will be compiled in embedded systems, and even Linux has had at times issues with /dev/urandom on Raspberry Pis.
[3] I would also like to see uint8_t, int8_t, uint16_t, int16_t, uint32_t, int32_t, uint64_t, and int64_t mandated. They exist in C99, but aren’t mandated, even though every real world compiler from this century supports all of the above types. Yes, I know about _BitInt(8/16/32/64/128/etc.) but a compiler from 2004—and yes I still use one to make win32 binaries—doesn’t support these new C23 datatypes.
We need to learn to use assembly and call it from C. C was never meant to be the end of programming. Reducing undefined behavior is the least the standards body can do. Fil-C is also an extremely useful tool.
The real problem I see is the inability of independent compiler writers upgrading to the latest standard. I think this is also an issue that the standards body should pay attention. Help implementers.
The concept of undefined behaviour specific to C/C++ has always seemed batshit insane to me, and I'm yet to read anything about it that has made it seem any less so.
i've seen static analyzers catch many ub patterns, but guaranteeing zero ub needs whole‑program analysis that blows up compile time and still produces false positives that drown developers
Make signed overflow defined please.
[dead]
[dead]
It’s cool that this mentions Fil-C but it also undersells it. TFA also undersells CHERI. Fil-C doesn’t just “find a lot of temporal-safety” bugs. It closes off memory safety bugs (special and temporal) for exploit writers and ascribes a tight semantics to the whole language. CHERI makes some different trade offs but also gives a tight enough semantics that memory safety exploits aren’t going to work. Both CHERI and Fil-C are more comprehensive than Rust, since they attack the problem at the ABI level (and so you don’t get the problem that the protection only applies to the parts that were rewritten in the safe subset of a new language). Rust could be claimed to be better in that its compile time, but that doesn’t make a significant difference if you’re worried about the definedness of semantics or exploitability.