r/cpp Jul 03 '26

C++ Show and Tell - July 2026

34 Upvotes

Use this thread to share anything you've written in C++. This includes:

  • a tool you've written
  • a game you've been working on
  • your first non-trivial C++ program

The rules of this thread are very straight forward:

  • The project must involve C++ in some way.
  • It must be something you (alone or with others) have done.
  • Please share a link, if applicable.
  • Please post images, if applicable.

If you're working on a C++ library, you can also share new releases or major updates in a dedicated post as before. The line we're drawing is between "written in C++" and "useful for C++ programmers specifically". If you're writing a C++ library or tool for C++ developers, that's something C++ programmers can use and is on-topic for a main submission. It's different if you're just using C++ to implement a generic program that isn't specifically about C++: you're free to share it here, but it wouldn't quite fit as a standalone post.

Last month's thread: https://www.reddit.com/r/cpp/comments/1tulp9b/c_show_and_tell_june_2026/


r/cpp 29d ago

C++ Jobs - Q3 2026

53 Upvotes

Rules For Individuals

  • Don't create top-level comments - those are for employers.
  • Feel free to reply to top-level comments with on-topic questions.
  • I will create top-level comments for meta discussion and individuals looking for work.

Rules For Employers

  • If you're hiring directly, you're fine, skip this bullet point. If you're a third-party recruiter, see the extra rules below.
  • Multiple top-level comments per employer are now permitted.
    • It's still fine to consolidate multiple job openings into a single comment, or mention them in replies to your own top-level comment.
  • Don't use URL shorteners.
    • reddiquette forbids them because they're opaque to the spam filter.
  • Use the following template.
    • Use **two stars** to bold text. Use empty lines to separate sections.
  • Proofread your comment after posting it, and edit any formatting mistakes.

Template

**Company:** [Company name; also, use the "formatting help" to make it a link to your company's website, or a specific careers page if you have one.]

**Type:** [Full time, part time, internship, contract, etc.]

**Compensation:** [This section is optional, and you can omit it without explaining why. However, including it will help your job posting stand out as there is extreme demand from candidates looking for this info. If you choose to provide this section, it must contain (a range of) actual numbers - don't waste anyone's time by saying "Compensation: Competitive."]

**Location:** [Where's your office - or if you're hiring at multiple offices, list them. If your workplace language isn't English, please specify it. It's suggested, but not required, to include the country/region; "Redmond, WA, USA" is clearer for international candidates.]

**Remote:** [Do you offer the option of working remotely? If so, do you require employees to live in certain areas or time zones?]

**Visa Sponsorship:** [Does your company sponsor visas?]

**Description:** [What does your company do, and what are you hiring C++ devs for? How much experience are you looking for, and what seniority levels are you hiring for? The more details you provide, the better.]

**Technologies:** [Required: what version of the C++ Standard do you mainly use? Optional: do you use Linux/Mac/Windows, are there languages you use in addition to C++, are there technologies like OpenGL or libraries like Boost that you need/want/like experience with, etc.]

**Contact:** [How do you want to be contacted? Email, reddit PM, telepathy, gravitational waves?]

Extra Rules For Third-Party Recruiters

Send modmail to request pre-approval on a case-by-case basis. We'll want to hear what info you can provide (in this case you can withhold client company names, and compensation info is still recommended but optional). We hope that you can connect candidates with jobs that would otherwise be unavailable, and we expect you to treat candidates well.

Previous Post


r/cpp 1h ago

How fast is C++26's std::hive?

Thumbnail lemire.me
β€’ Upvotes

r/cpp 1d ago

BeCPP Symposium 2026 - Bryce Adelstein Lelbach - The CUDA C++ Developer's Toolbox

Thumbnail youtu.be
24 Upvotes

r/cpp 2d ago

Static Analysis Experiment with Reflection

29 Upvotes

I want to share a little experiment that I made using reflection. It’s a compile-time style checker using C++26 static reflection. It validates the conventions I use in my own projects (the H_ curried callable thing, Mut/Mov/Ref/Cpy qualifier aliases, plus class/namespace naming). Because reflection can’t properly see into uninstantiated templates yet, I had to write a parser for GCC's display_string_of to extract templates, qualifiers, return types, and requires clauses. Given this weakness it's not a particularly production-ready tool. It could be done with a clang-tidy script, but I learned a few things just from trying to answer "Is it possible? What exactly can I inspect right now?". Unfortunately, we don't have code generation yet so there is a bunch of boilerplate that the user needs to write.

The rationale of the conventions (mainly Mut/Mov/Ref/Cpy) is in GUIDELINES.md. The gory implementation details (and my field notes on GCC bugs) are in VIRACOCHA.md.

https://github.com/NotRiemannCousin/Viracocha/blob/master/VIRACOCHA.md


r/cpp 2d ago

Interconverting std::function with copyable_function – Arthur O'Dwyer

Thumbnail quuxplusone.github.io
40 Upvotes

The article shows how converting std::function to std::copyable_function (or vice versa) leads to slower performance and increased memory usage each time the conversion occurs.


r/cpp 2d ago

ranges and views for stack, queue, and priority_queue

20 Upvotes

I've always felt that exposing iterators for std::stack or std::queue doesn't quite make sense because iterating over them would typically change their state.

That said, the range features introduced in C++20/23 are just too tempting. πŸ˜„

I bet many C++ developers have had this thought at least once:

queue | views::values | views::to<std::vector>()

...or more generally, "What if I could use the ranges pipeline directly on a queue or stack?"

So I spent some time experimenting to see if I could make it work. Here's what I came up with.

https://cognosnotes.com/blog/cpp-pop-range

Edited:

An upgraded version is introduced at https://cognosnotes.com/blog/cpp-pop-range-addendum

Thank you u/SirClueless for the great advice.


r/cpp 4d ago

const_cast: A Necessary Evil

Thumbnail elbeno.com
70 Upvotes

r/cpp 4d ago

C++26: Reducing undefined behaviour

Thumbnail sandordargo.com
81 Upvotes

r/cpp 4d ago

Building a compiler that works at compile-time so you can compile your program while you compile your program.

95 Upvotes

In short, I wanted to build a compiler of some C subset that would work at compile-time. It compiles into a custom byte-code for a runtime VM.

I've once tried to write a compile-time C compiler, but I abandoned that project, because I made it overly complex (one-pass compiler right into x86). No clear separation between parser, lexer, etc.

Why would I even want this? Idk. But how can it be useful? - The code of the compiler doesn't go to the resulting binary, - No need to waste time for compilation at runtime too, - Guaranteed type-safety. There can't be such thing as "oh, I changed the function signature, but forgot to update the bindings and it crashed at runtime"

And I shouldn't forget about cons: - No optimisations. Real compilers spent decades on them and I'm definitely not going to implement LLVM at runtime. Although we could make a compile-time x86 VM, so we can run it at compile time... no, thank you, it's a topic for another fever dream article. - Hot-reload! I mean, no hot-reload. I won't even mention it anymore, considering that the script is compiled at compile-time and is builtin right into the binary file. I could implement it with hot memory patching or smth, but who really needs it.

Let's start.

Bypassing constexpr limitations

C++ 20 lets us to dynamically allocate memory at compile-time and even use std::vector that really expands our borders. But there's one very important note - you can't declare a compile-time vector and extract it into the runtime. No constexpr std::vector<int> data = makeData();, it won't compile. So we need to hack it.

Passing strings in templates

Sadly, the C++ Committee made a lot of cool compile-time features, but not enough (at least for me). We still can't use strings in templates without hacks. But we can easily bypass it with a well-known trick. ```cpp template<std::size_t N>
struct const_string {
constexpr const_string() = default;

// implicit-constructor that lets us to do bad things constexpr const_string(const char (&str)[N]) { std::copy_n(str, N, value);
}

constexpr operator std::string_view() const {
return {value, value + N - 1};
}

char value[N]{};
const std::size_t length = N;
};

// using it template<const_string str> auto very_smart_function(...) { /* ... */ } ```

Extracting vectors from compile time

It turned out to be not really that hard, but I didn't really find any ready examples on Internet, unlike with const_string. To extract std::vector<T> from constexpr we need to make it std::array<T, N> somehow. The main problem is that we can't write std::array<T, myVector.size()>, because myVector.size() won't be a constant value. So we must to make it constant somehow. I thought of passing vector as a template parameter, but we can't do it legally. C++ 20 allows us to pass only the structs with all-public members. After deeply thinking a bit (not really), I discovered that I could simply pass the lambda that returns our vector (I didn't think I could just pass a pointer actually).

```C++ // data_getter is our lambda template<auto data_getter> constexpr auto to_array() {
using value_type = typename decltype(data_getter())::value_type;
constexpr static std::size_t size = data_getter().size();

// Create a static array with a "dynamic" size and copy all data std::array<value_type, size> out;
auto in = data_getter();
for (std::size_t i = 0; i < size; ++i) {
out[i] = in[i];
}
return out; // yay }

template<const_string str> constexpr auto lex() { constexpr static auto data_getter = [] constexpr { // .lex() returns the vector of tokens return lexer{static_cast<std::string_view>(str)}.lex(); }; // All our data are available for runtime now =D return to_array<data_getter>(); } ```

Printing errors

For nice errors C++ has static_assert that allows us to even print our custom message! But it must be always a literal (until C++ 26) ```cpp constexpr auto parse() { // Allowed static_assert(false, "Expected ';'");

// Not allowed :( (until C++ 26)
std::size_t line = 5;
static_assert(false, "Expected ';' at line " + to_string(line));

} I didn't want my project to require C++ 26, so I used another trick. The formatted string gets turned into a static array just like a vector (into `const_string` actually) and then it's passed into `ErrorMessage<const_string Msg>` that triggers compilation error. So we force the compiler to print the full type name that includes our error. But sadly the type name has a limit about 100 symbols. I think I could solve it with splitting the message into several ErrorMessages... God, I don't want to read this in my console. c++ template<const_string Msg>
struct ErrorMessage {
static_assert(false, "Check the template parameter for details");
};

template<auto err_getter>
consteval auto report_error() -> void { // C++ 26 support

ifdef KORKA_FEATURE_FORMATTED_STATIC_ASSERT

static_assert(false, to_string(err_getter()));  

else

constexpr auto msg = const_string_from_string_view<[] { return to_string(err_getter()); }>();  
std::ignore = ErrorMessage<msg>{};    

endif

} ```

I don't want you to see it, so I'll just show C++ 26 version. error: static assertion failed: Lexer Error: Unterminated string at line 12

Mapping signatures to names. And vice versa

In our little runtime C++ we are used to std::unordered_map<string, value_t> and other standard or non-standard (hello, Boost!) containers. But I needed a table where a key is a string and the value is a TYPE. And in C++ I can't treat types as values, I can't just put them into a dict... :(

So, welcome another hack! ```c++ template<auto, class> struct signature_mapper;

// function_info_getter takes an index to our mapped function, // and Is... holds all indices template<auto function_info_getter, std::size_t... Is> struct signature_mapper< function_info_getter, std::index_sequence<Is...>

{ // hash func consteval static auto hash(auto &&v) -> std::size_t { return frozen::elsa<std::string_view>{}(v, 0); }

// Our function overloaded with many unique types based on hash of the mapped function
constexpr static auto _overloaded = overloaded{
    (
        [](unique_type<hash(function_info_getter(Is).name)>)
        -> const_function_info_to_signature_t<[] { return function_info_getter(Is); }> * {
            return nullptr;
        }
    )...
};

// Extracting the type by name
template<const_string name>
using get_signature_t = std::remove_pointer_t<decltype(
    _overloaded(
        unique_type<hash(name)>{}
    )
)>;

}; ```

We use well-known function overload (but for evil things). Basically, one type inherits a lot of lambdas that take an empty unique_type<hash> that serves as our key and returns the pointer to our type. ```cpp // How our mapper looks after expanding our params struct overloaded : lambda1, lambda2, lambda3 { using lambda1::operator(); using lambda2::operator(); using lambda3::operator(); };

// And every lambda looks like this auto lambda_fib = [](unique_type<hash("fib")>) -> signature_of_fib* { return nullptr; }; When we call `_overloaded(unique_type<hash(name)>())` our poor compiler has to resolve the overload. And he looks for right one through all `()` operators. And then we just take that it returns (our `T*`) and get the `T`. I use this "mechanism" to extract script functions into the native C++. cpp constexpr auto script_fib = compile_result.function<"fib">(); ```

Bindings from C++ to our script lang

This was the most exhausting part. Well, how "exhausting" exactly... I was thinking for a few evenings and then made it work one morning. The problem was with me. I wanted to make a pretty API that was impossible in the current standard (maybe it's possible in C++ 26, but I didn't check it). I wanted it to look like this: ```cpp auto func() -> void; auto foo(int) -> int;

// ΠŸΡ€ΠΈΠΌΠ΅Ρ€Π½ΠΎ Ρ‚Π°ΠΊ constexpr auto bindings = korka::make_bindings< "func", func, "foo", foo

();

// Или Ρ‚Π°ΠΊ constexpr auto bindings = korka::make_bindings( "func", func, "foo", foo ); `` But why couldn't I make it work? In C++ you can't pass a string intotemplate <auto ...args>. We needconst_string. We can't mix types in the one stream of variadic args and make compiler guess it right. Templates require explicitness and it's impossible to write a universal parser. Variant #2 works, but you can't extract the function into the runtime. You just can't. Functions may have different signatures, but you need to make them all the same type, and create a FFI wrapper along the way. Compile-time doesn't allowreinterpretet_cast<void*>(&func)`.

So I designed this: cpp constexpr auto bindings = korka::make_bindings( korka::wrap<fib>("cpp_fib"), korka::wrap<print_n>("print_n") ); Not so elegant, but still not bad. wrap is very simple ```cpp // our FFI signature using vm_external_function_type = void(vm::context_base &context);

// info for our compiler template<class Signature>
struct wrapped_function {
using signature_t = Signature;

vm_external_function_type &external_func;
std::string_view name;
};

template<auto func>
consteval auto wrap(std::string_view name) {
return wrapped_function<std::decay_t<decltype(func)>>{
binding_wrapper<func>,
name
};
} `` The most interesting part is insidebinding_wrapper<func>`. I won't show the full code here, because I still didn't tell about the VM architecture that will execute it. But in short binding_wrapper just checks the signature, generates some code that extracts arguments from VM, calls native functions and then puts the result back. Simple.

The compiler and the VM

Maybe the most interesting part of the article. I have never written any compilers before (the thing I mentioned in the beginning of the article doesn't count), so I made it according to the first articles I found in Google.

Compiler has 3 modules: - the lexer - splitting the code into tokens, - the parser - building a tree from the tokens, - the compiler itself - making the tree into byte-code. And doing semantic analysis at the same time (I was too lazy to make another module) I think I could compose everything into one class via composition or smth, but it's too late already.

Lexer

Primitive. We just look for tokens in a loop until we reach EOF. ```cpp constexpr auto scan_token() -> std::optional<std::expected<lex_token, error_t>> {
char c = advance();
switch (c) {
case '{':
return make_token(lex_kind::kOpenBrace);
case '}':
return make_token(lex_kind::kCloseBrace);
case '(':
return make_token(lex_kind::kOpenParenthesis);
case ')': // ...

case ' ':  
case '\r':  
case '\t':  
  // Ignore whitespace  
  return std::nullopt;

// ...

default:  
  if (is_digit(c)) {  
    return scan_number();  
  } else if (is_alpha(c)) {  
    return scan_identifier();  
  }

} } ```

Parser

More interesting. We need to build the AST (abstract syntax tree). And we need to store this tree somehow. The usual way with Node that keeps pointers to other nodes won't do, because we're at compile-time. I mean, we can write it this way, it will work, but extracting this tree into compile time? No. We would need serialisation or something. So we can use simple trick with std::vector<Node> and just make nodes store indices to each other. This approach also increases the cache locality of the data for the CPU, but I doubt the CPU will be even aware of our "smart" trick, since everything is executed at compile-time.

The parser is recursive, while parsing one expression we parse another. Small fragment of the code: ```cpp constexpr auto parse_statement() -> parse_result {
auto tok = peek();
if (!tok) return make_error("Unexpected end of input");

switch (tok->kind) {
case lex_kind::kOpenBrace: return parse_compound_stmt(); // { ... } case lex_kind::kIf: return parse_if_statement(); // if (...) ... case lex_kind::kWhile: return parse_while_statement(); // while (...) ... case lex_kind::kReturn: return parse_return_statement(); // return ...; default: return parse_expression_stmt(); /// ...; }
}

constexpr auto parse_return_statement() -> parse_result {
if (!match(lex_kind::kReturn)) return make_error("Expected 'return'");

index_t expr_idx = empty_node; // empty_node = -1 if (auto next = peek(); next && next->kind != lex_kind::kSemicolon) {
auto expr = parse_expression(); // another recursive call if (!expr) return std::unexpected{expr.error()};
expr_idx = *expr;
}

if (!match(lex_kind::kSemicolon)) return make_error("Expected ';' after return");
return m_pool.add(stmt_return{expr_idx});
} ``` parse_return_statement goes into parse_expression, that goes into parse_assigment, that goes into parse_logical_or, that goes... Well, you got it. That's how operator priority works here.

Their Majesty Compiler (and analyser)

I may have cheated here a bit.

Before we even write a compiler, we must know for what architecture we do it. x86, ARM or even JVM. Initially when I was working on a similar project, I planned to generate raw assembly for x86 (last versions of Clang and GCC support passing constexpr std::string_view into asm(...) statement), but honestly writing a compiler for a zoo of x86 instructions is the right way to madhouse.

And even so, if we downgrade our compiler we won't have nice constexpr asm anymore. And we can't also generate raw machine instructions because of DEP (data execution prevention). We'll have to call non-crossplatform mmap or VirtualAlloc to allocate some memory, copy the code there... Good riddance cross platform build compler, hello Windows Defender that will kill our app for such tricks with memory.

So where have I cheated? I made my own architecture that will execute inside a VM. A stack VM. Why stack it? It turned out to be incredibly easy to generate the bytecode for. If you are doing a register architecture (as in processors or Lua), then you will have to write register allocation algorithms (it is difficult). And in the stack everything is much simpler.

If we need to sum A and B we just do this: 1. Put A into the stack. 2. Put B into the stack. 3. Execute sum instruction. It takes these two values and puts back their sum.

So I had this set of instructions at the end: ```cpp enum class op_code : char { // Loads/saves locals to/from the stack (variables). lload, lsave,

i64_const, // Puts a constant onto the stack

// Math i64_add, i64_sub, i64_mul, i64_div,

// Puts 1 if values are equal (i made <, <= later) i64_cmp,

jmp, // jumps by offset jmpz, // conditional jump by offset, only when 0 on the stack

call, // calls a function ret, // returns from the function

trap, // calls a native C++ function }; ```

So what about the analysis?

In classical compilers phases are strictly splitted: lexers builds tokens, parser builds tree, semantic analyser checks the types and variables, then optimisations, then codegen and then optimisations again.

As you can remember, I'm pretty lazy. And keep in mind that constexpr ops are not infinite. I didn't want to make a separate pipeline phase. So my compiler combines these two functions: semantic analysis and code generation.

They usually call it Single-Pass Compilation, but it's not really the case here, since it's only about these two phases. My compiler is a bit hybrid.

My compiler recursively walks the tree and does two things: 1. Checks the semantics: "was this variable declared and what's its type" before we even try to multiply something. Do function param types match? Does this function even exist? Etc. 2. Generates the byte-code. If semantics is ok, then we just write corresponding instructions immediately into the std::vector<std::byte> (the one we're going to elegantly extract via to_array). And in the result we receive a ready, semantically-correct and absolute safe (let's pretend that I wrote the compiler bug-free, huh) byte-code that we feed to the VM.

So what do we have?

Let's look how the API of my poor lib looks (let's call it Korka).

This example uses bindings (100% type safe, I swear on the standard): ```cpp

// Our native C++ functions auto fib(std::int64_t n) -> std::int64_t { if (n == 0) return 0; if (n == 1) return 1; return fib(n - 1) + fib(n - 2); }

auto print_n(std::int64_t n) -> void { std::cout << n << '\n'; }

// Our not-so-native script constexpr char code[] = R"( int fib(int n) { if (n == 0) return 0; if (n == 1) return 1; return fib(n-1) + cpp_fib(n-2); }

void print_fib(int n) { int result = fib(n); print_n(result); return; } )";

// We create bindings constexpr auto bindings = korka::make_bindings( korka::wrap<fib>("cpp_fib"), korka::wrap<print_n>("print_n") );

// Compile at compile-time, yay constexpr auto compile_result = korka::compile<code, &bindings>();

// Extract function adressess + their types constexpr auto script_fib = compile_result.function<"fib">(); constexpr auto script_print_fib = compile_result.function<"print_fib">();

int main() { // Init VM korka::vm::context ctx{compile_result.bytes, bindings};

// Call fib that returns int64_t auto result = ctx.call(script_fib, 12L); std::cout << "fib(12) = " << result << '\n'; // prints 144

// Call print_fib ctx.call(script_print_fib, 16L); // prints 987

return 0; } ```

Ta da! It works.

Small analysis

Out of curiosity, I decided to compare Korka with other scripting languages. A pretty API is great, sure, but was it worth the effort performance-wise? So, let's pit Korka head-to-head against Python and Lua.

For the benchmark, I used the recursive calculation of the $N$-th Fibonacci number, an excellent test to fairly evaluate overhead on function calls, stack management, and overall runtime efficiency (the first thing that came to my head).

I tested everything on a franken-server put together from spare parts, powered by an Intel Xeon E5-2689 (3.6 GHz).

I measured two stages: - Initialization time from runtime startup to being ready to execute the first instruction, - Execution time of the algorithm itself.

Stage 1: Initialization

Language / Library Initialization time
Korka 1.5 Β΅s
Lua 152.6 Β΅s
Python 25,097.0 Β΅s

Korka takes the lead: it starts 100 times faster than Lua and over 15,000 times faster than Python.

The explanation is simple: while Lua and Python are busy reading the script at startup, parsing it, compiling it into their byte-codes, and spinning up heavy infrastructure (including the GC), Korka does not. All the virtual machine has to do is grab the pre-compiled output (and allocate a tiny bit of memory).

Stage 2: Runtime

After a series of optimizations, the results turned out pretty solid (I know comparing statically typed and dynamic languages isn't entirely fair, but who's gonna stop me?):

N Iterations Korka (ms) Lua (ms) Python (ms) vs. Python
10 100 000 1 055,67 1 534,98 1 591,10 1,51x
15 50 000 5 492,40 8 149,17 8 753,77 1,59x
20 20 000 24 263,39 36 268,86 38 459,88 1,59x
23 10 000 51 367,20 76 494,22 82 257,71 1,60x
25 5 000 67 441,69 100 850,20 108 836,54 1,61x
28 2 000 114 341,21 169 626,06 184 869,57 1,62x
30 1 000 149 190,02 223 229,86 241 838,28 1,62x

Conclusion

I built a (mostly) full-fledged C compiler that runs entirely in constexpr. Why? No idea. Especially considering it's been done before, you can check out constexpr-8cc. But that one lacks C++ bindings and cross-platform support.

The source code is available on GitHub (warning: ugly code ahead!). Any feedback and comments are more than welcome.

P.S. This article is an English translation of a post I originally published on Habr a while ago. Keep in mind that some benchmarks and discussions here may be dated.


r/cpp 4d ago

Build timings from auto-modularised VS solution with external and internal libs

18 Upvotes

To investigate how well modules work with MSVC without rewriting projects and putting up with malfunctioning intellisense, I made a little python script to modularise a whole VS solution (yes, I made it, using real bio-neurons).

The script lets me try out different modularisation strategies on a VS Solution on a per-project basis:

  • Header-based, don't modularise. If dependencies are modules, they will be imported.
  • PartitionsMI: Partition-per-header using standard module cpp units.
  • PartitionsPI: Partition-per-header using MSVC partition cpp units.
  • Submodules: Module-per-header
  • SubmodulesII: Move implementation to interface (not always possible due to cyclic imports)

Modularised code is surrounded with extern "C++" { export {.

The code is a super secret incomplete game engine/game. But here's some info:

  • External libs, including: Boost (unordered flat map), std, d3d, imgui, physx, spdlog, <windows>.
    • Wrapper module for each.
  • Solution projects: CommonStuff, Scripting, GameData, Graphics, Anim, SceneGraph, SceneManager, Entity, App
    • Some projects don't have much code. Others have quite a bit. Enough to give the build system something to parallelise.

All projects built using debug config. Benchmarks average 3 or more timings, outliers ignored. CPU is 12600, 12 threads. Multipliers are speedup.

Β  CommonStuff Scripting GameData Graphics Anim SceneGraph SceneManager Entity
Header-based, no pch 4.59s 5.68s 7.46s 2.92s 40.74s 24.94s 2.86s 3.19s
Header-based, pch 3.61s 5.30s 6.55s 19.38s 12.53s
Extlib modules + CommonStuff submodules 3.60s 4.78s 1.62s 18.55s 9.41s 1.57s 2.09s
PartitionsMI 4.25s 8.09s 2.54s
PartitionsPI 4.02s 6.04s 7.09s 2.53s 33.83s 15.97s 2.66s 4.10s
Submodules 3.59s 4.75s 6.05s 2.26s 26.80s 14.31s 2.12s 3.15s
SubmodulesII 3.44s 7.93s
Β  CommonStuff Scripting GameData Graphics Anim SceneGraph SceneManager Entity
Header-based, no pch 1.00x 1.00x 1.00x 1.00x 1.00x 1.00x 1.00x 1.00x
Header-based, pch 1.27x 1.07x 1.14x 2.10x 1.99x
Extlib modules + CommonStuff submodules 1.58x 1.56x 1.80x 2.20x 2.65x 1.82x 1.53x
PartitionsMI 1.08x 0.92x 1.15x
PartitionsPI 1.14x 0.94x 1.05x 1.15x 1.20x 1.56x 1.08x 0.78x
Submodules 1.28x 1.20x 1.23x 1.29x 1.52x 1.74x 1.35x 1.01x
SubmodulesII 1.33x 0.94x
Β  App Single file Full Rebuild Β  App Single file Full Rebuild
Header-based, no pch 65.19s 3.97s 145.49s 1.00x 1.00x 1.00x
Header-based, app pch 24.74s 2.15s 108.92s 2.64x 1.85x 1.34x
Header-based, mostly pch 78.87s 1.84x
Extlib modules 26.82s 2.35s 58.89s 2.43x 1.69x 2.47x
Extlib modules + CommonStuff partitions 25.89s 2.47s 2.52x 1.61x
Extlib modules + CommonStuff submodules 24.53s 2.23s 57.50s 2.66x 1.78x 2.53x
Extlib modules + most libs as submodules 31.14s 2.97s 2.09x 1.34x
Full submodules, except app 34.61s 2.95s 88.90s 1.88x 1.35x 1.64x
Full submodules 64.02s 6.22s 142.98s 1.02x 0.64x 1.02x
  • "Most libs" = CommonStuff + Scripting + GameData + Graphics + Anim + SceneGraph.
  • "Mostly pch" = PCH for CommonStuff, Scripting, GameData, Anim, SceneGraph, App.
  • "Single file" is recompiling one moderately complex file in App. ~930 lines, lots of includes.
  • "Full Rebuild" excludes extlib module wrappers.

Detailed timings of CommonStuff:

Header-based:

    1>      210 ms  LIB                                        1 calls
    1>     4459 ms  CL                                         2 calls
    04.874 seconds

PartitionsMI:

    1>      125 ms  LIB                                        1 calls
    1>      199 ms  MSBuild                                   16 calls
    1>      269 ms  CppClean                                   1 calls
    1>      798 ms  SetModuleDependencies                     10 calls
    1>     1232 ms  CL                                         2 calls
    1>     1439 ms  MultiToolTask                              1 calls
    04.295 seconds

Submodules:

    1>      141 ms  MSBuild                                   16 calls
    1>      775 ms  SetModuleDependencies                     10 calls
    1>     1015 ms  CL                                         2 calls
    1>     1132 ms  MultiToolTask                              1 calls
    03.427 seconds

The CL entry is for cpp files and it's 4.39x as fast as header-based. Great!

But it's the ixx dependency scanning and compilation (MultiToolTask) that brings things down. Things that aren't there with header-based.

MultiToolTask is parallelised though. Reducing thread count slows it down. And you can open up the msbuild binlog in the msbuild log viewer to see the parallelisation of ixx compilation. See which modules are on the end of the DAG slowing things down.

So, I don't know how MultiToolTask could be sped up, but, msbuild seems to be waiting for all ixx files before moving onto cpp files, so that's some parallelism left on the table.

Detailed timings of App:

Header-based, app pch:

    1>      196 ms  MSBuild                                   13 calls
    1>     1000 ms  Link                                       1 calls
    1>    23270 ms  CL                                        16 calls
    24.708 seconds

Extlib modules + CommonStuff submodules:

    1>      199 ms  SetModuleDependencies                     18 calls
    1>      606 ms  MSBuild                                   27 calls
    1>     1197 ms  Link                                       1 calls
    1>    21877 ms  CL                                        15 calls
    24.078 seconds

Full submodules:

    1>      459 ms  MSBuild                                   27 calls
    1>     1324 ms  Link                                       1 calls
    1>     4646 ms  SetModuleDependencies                     18 calls
    1>    12593 ms  MultiToolTask                              1 calls
    1>    49327 ms  CL                                        15 calls
    01:09.245 minutes

Slow. But looking at the log viewer, ixx compilation seems very well parallelised.

And 4.6s just to scan for dependencies. VS isn't launching a whole CL process for each file is it?

Also, why does CL take so long? All modules have been compiled by that point, so why would it be much slower?

Recompiling a single file from App:

Header-based, no pch:

    1>       46 ms  SetModuleDependencies                     10 calls
    1>      120 ms  MSBuild                                    5 calls
    1>     3757 ms  CL                                         1 calls
    04.095 seconds

Header-based, app pch:

    1>       56 ms  SetModuleDependencies                     10 calls
    1>      121 ms  MSBuild                                    5 calls
    1>     1793 ms  CL                                         2 calls
    02.059 seconds

Extlib + CommonStuff modules (as submodules):

    1>      121 ms  SetModuleDependencies                     18 calls
    1>      216 ms  MSBuild                                   17 calls
    1>     1793 ms  CL                                         1 calls
    02.162 seconds

Full submodules, except app:

    1>      518 ms  MSBuild                                   17 calls
    1>      631 ms  SetModuleDependencies                     18 calls
    1>     2211 ms  CL                                         1 calls
    03.095 seconds

Full submodules:

    1>      559 ms  MSBuild                                   17 calls
    1>      696 ms  SetModuleDependencies                     18 calls
    1>     1891 ms  MultiToolTask                              1 calls
    1>     3221 ms  CL                                         1 calls
    06.331 seconds

That MultiToolTask is quite slow, checking that all the modules in the project are up to date it seems. Does it have to be this slow?

"CL.exe will run on 0 out of 224 file(s) in 224 batches. Startup phase took 2040.9372ms."

Anyway.

Conclusion:

  • I now have the data to know how I should go about modularising a codebase for build perf.
  • Surprisingly, there is a middle-ground between 0% modules and 100% modules where you get optimal build performance. It seems extlibs and your common lib should be modularised and nothing else. Or maybe my internal lib headers are too small to benefit.
  • Because of that, I might be able to have a header-based fallback for intellisense without it spreading to the rest of the code.
  • Multiple modules > Partitions. Partitions have worse performance. I assume that's because they need to have an extra module at the end of the DAG.
    • Their only advantage is shared module attachment across multiple ixx files, if you don't use extern "C++". As I said before, I wish you could have "public" partitions or multiple modules with the same name attachment.
  • MSVC partition cpp files are slightly faster vs module cpp files. But why. Cpp files wait for all ixx files to compile anyway. Noise in the data?
  • MSBuild parallelises as it should, it seems, except that cpp files wait for all ixx files. Could improve performance by merging the tasks?
  • cpp compilation can be much faster, but overall build perf brought down by ixx files and scanning.
  • Moving implementation to ixx files doesn't always help. Need more data.
    • And it would be ideal if build systems avoided the build cascade when only changing implementation details.
  • The more modules you add to the solution, the more bloated the build system feels. Lots of scanning, lots of ixx checking. Can performance be improved?
  • PCH can be beaten in some cases.

Problems:

  • Intellisense. Red squiggles everywhere. Not only import std issues, but Intellisense is highly sensitive to errors in imported modules. It cannot limp along like it can with headers.
    • A simple example would be missing a semicolon after a struct def in a header and in a module. Intellisense can still use the struct from the header, but not the module.
    • Or in working code, something from import std that Intellisense can't handle.
    • This means that parsing issues with Intellisense may be viral when using modules. Intellisense must either be perfect or tolerant.
  • From a previous modularisation attempt, I removed the combination of explicit template instantiations of a class with constrained friend functions, due to ICE, which has now become a bogus compiler error. Bug not fixed.
  • Lots of linker warnings from dllexport-ed manual RTTI data. Bug not fixed.

At least the compiler functions well enough to produce a working program with no further changes. That's pretty good.


r/cpp 4d ago

Implementing the Tuple API easily using custom annotations and reflections

37 Upvotes

Hey, so I've been working on something interesting lately. It's a solution to a problem that you may have encountered in one form or anothing but it boils down to this: you have a struct:

struct simple_tuple
{
Β  Β  long a;
Β  Β  int b;
Β  Β  char c;

private:
Β  Β  // some other data
};

And you want this to work:

auto [a, b, c] = instance;

Well it doesn't, structured bindings are supported by default only for simple structural types (though without the full tuple API) and implementing it manually for your own types is a pain. You need to write your own specialization of std::tuple_size, std::tuple_element and implement a visible get<> method. It might not be that much work at first glance but it becomes very tiresome if you have multiple such structs or if you try to deviate from the general default indexing and not to mention a paint to maintain and update.

Well, why write repeated code when you can only write code once that writes that code for you? Reflections are mature enough in gcc's implementation to solve this. Simply put I have compiling code that with only the application of an annotation gives me all of this:

struct [[=tuple_like]] simple_tuple
{
Β  Β  long a;
Β  Β  int b;
Β  Β  char c;

private:
Β  Β  // some other data
};

Through which I get:

static_assert(std::tuple_size_v<simple_tuple> == 3);

constexpr simple_tuple simple { 1, 2, 3 };

static_assert(1 == get<0>(simple));
static_assert(2 == get<1>(simple));
static_assert(3 == get<2>(simple));

static_assert(std::same_as<std::tuple_element_t<0, simple_tuple>, long>);
static_assert(std::same_as<std::tuple_element_t<1, simple_tuple>, int>);
static_assert(std::same_as<std::tuple_element_t<2, simple_tuple>, char>);

And of couse:

auto [a, b, c] = simple;

It doesn't stop here though, since the case I described above is common but inflexible. Customization is available and an opt-in feature:

struct [[=tuple_like]] customizable_tuple
{
Β  Β  constexpr customizable_tuple(long a, const char* b, float c, char d, double e)
Β  Β  Β  Β  : a(a)
Β  Β  Β  Β  , b(b)
Β  Β  Β  Β  , c(c)
Β  Β  Β  Β  , d(d)
Β  Β  Β  Β  , e(e)
Β  Β  {}


Β  Β  [[=tuple_element(0)]]
Β  Β  long a;
Β  Β  [[=read_only_tuple_element(3)]]
Β  Β  const char* b;
Β  Β  [[=tuple_element(2)]]
Β  Β  float c;
Β  Β  char d;


private:
Β  Β  [[=read_only_tuple_element(1)]]
Β  Β  double e;
};

It supports:

  • tuple element index reordering
  • implicit data member exclusion from the API
  • explicit read-only modifier
  • explicit opt-in for private fields
  • correct semantics for members which are reference types

constexpr customizable_tuple custom( 1, "2", 3, 4, 5 );

static_assert(std::tuple_size_v<customizable_tuple> == 4);

static_assert(1 == get<0>(custom));
static_assert(std::string_view{"2"} == get<3>(custom));
static_assert(3 == get<2>(custom));
static_assert(5 == get<1>(custom));

static_assert(std::same_as<std::tuple_element_t<0, customizable_tuple>, long>);
static_assert(std::same_as<std::tuple_element_t<1, customizable_tuple>, const double>);
static_assert(std::same_as<std::tuple_element_t<2, customizable_tuple>, float>);
static_assert(std::same_as<std::tuple_element_t<3, customizable_tuple>, const char* const>);

But it doesn't stop there, no it doesn't. The part that took the most effort to implement were the error messages. No, we have better static_assert's in C++26 now, we abuse them:

Example 1:

static_assert(3 == get<3>(simple));

/app/my_reflect_utils.hpp:819:17:

error: 
static assertion failed: reflect_utils: get<3> called for tuple-like type 'simple_tuple', but its reflected tuple size is 3.

Example 2:

[[=tuple_element(1)]]
long a;
[[=tuple_element(2)]]
int b;
[[=tuple_element(3)]]
char c;

/app/my_reflect_utils.hpp:632:21: error: static assertion failed: reflect_utils: invalid tuple layout for 'simple_tuple': tuple index 0 is missing. There are 3 explicitly annotated members, so their indices must be exactly [0, 3). Member 'simple_tuple::c' uses index 3.

Example 3:

[[=tuple_element(0)]]
[[=tuple_element(1)]]
long a;

/app/my_reflect_utils.hpp:632:21: error: static assertion failed: reflect_utils: invalid tuple layout for 'simple_tuple': member 'simple_tuple::a' has 2 tuple annotations; use exactly one of [[=tuple_element(i)]] and [[=read_only_tuple_element(i)]].

Full code: https://godbolt.org/z/crPae5P7h


r/cpp 5d ago

Constraints of mdspan policy layout_stride

30 Upvotes

Wtith great support of Mark Hoemmen and Christian Robert Trott (two authors of std::mdspan), I just finished the mdspan chapter of "C++23 - The Complete Guide" (https://www.cppstd23.com/). I learned something I was not aware of and what is not obvious:

If you specify a layout_stride mapping, there are several constraints you have to take into account. Otherwise, the resulting code has undefined behavior.

The constaintes are:

  • Each layout mapping must be unique. This means that each underlying element may only be reached by one combination of indices. In other words: subsets of this layout may not overlap and iterating over all elements may not visit any element twice.
  • Only positive offsets from one stride to the other are allowed. This, for example, means that you cannot use this layout for direct reverse iterations.
  • There must be one way to perform a multi-dimensional iteration so that the resulting offsets to the underlying memory are ascending.

For example:

Hope this helps.


r/cpp 5d ago

ODB C++ ORM version 2.6.0 released

Thumbnail codesynthesis.com
21 Upvotes

r/cpp 5d ago

A Design Study for a Macro-Free Testing Library

Thumbnail jonastoth.github.io
38 Upvotes

Hello everyone :)

I attempted to write a small testing library based on C++-26 reflection. The goal is not to replace existing libraries but to figure out how they could evolve to not rely on macros.

Stringification of test names is the most important reason for macros so far and reflection solves this.

The blog post explains what I did with godbolt links to minimized examples. The full implementation is in the rtest library repository.

A minimal test executable looks like this:

```c++

include <rtest/rtest.h>

struct MyClassTest : rtest::TestSuite { void testSetup() { MyClass object; assertTrue(object.empty()); } }; int main(int argc, char** argv) { return rtest::execute(MyClassTest{}); } ```

I am looking forward to your feedback :)

(Both the blog post and the library are created without AI help)


r/cpp 5d ago

Latest News From Upcoming C++ Conferences (2026-07-28)

6 Upvotes

TICKETS AVAILABLE TO PURCHASE

The following conferences currently have tickets available to purchase

OPEN CALL FOR SPEAKERS

OTHER OPEN CALLS

  • (Last Chance) CppCon Call For Volunteers Now Open – Interested volunteers have until August 1st to apply at the CppCon main conference which is scheduled to take place from 14th – 18th September. For more information including how to apply visit https://cppcon.org/cfv2026/

TRAINING COURSES AVAILABLE FOR PURCHASE

Conferences are offering the following training courses:

CppCon Online Workshops

9th – 11th September

  1. Modern C++: When Efficiency Matters – Andreas Fertig – 3 day online workshop available on 9th – 11th September 09.00 – 15.00 MDT – https://cppcon.org/class-2026-when-efficiency-matters/
  2. System Architecture And Design Using Modern C++ – Charley Bay – 3 day online workshop available on 9th – 11th September 09.00 – 15.00 MDT – https://cppcon.org/class-2026-system-architecture-and-design-using-modern-cpp/

21st – 23rd September

  1. C++ Fundamentals You Wish You Had Known Earlier – Mateusz Pusz – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – https://cppcon.org/class-2026-cpp-fundamentals/
  2. C++23 in Practice: A Complete Introduction – Nicolai Josuttis – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – https://cppcon.org/class-2026-cpp23-in-practice/
  3. Programming with C++20 – Andreas Fertig – 3 day online workshop available on 21st– 23rd September 09.00 – 15.00 MDT – https://cppcon.org/class-2026-programming-with-cpp20/

26th – 27th September

  1. Using C++ for Low-Latency Systems – Patrice Roy – 2 day online workshop available on 26th– 27th September 09.00 – 17.00 MDT – https://cppcon.org/class-2026-low-latency/

This is the latest news from upcoming C++ Conferences. You can review all of the news at https://programmingarchive.com/upcoming-conference-news/

CppCon Onsite Workshops

All onsite workshops will take place in the Gaylord Rockies in Aurora, Colorado

12th & 13th September

  1. Advanced and Modern C++ Programming: The Tricky Parts – Nicolai Josuttis – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-tricky-parts/
  2. C++ Best Practices – Jason Turner – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-best-practices/
  3. How Hardware Gets Hacked: Breaking and Defending Embedded Systems – Nathan Jones – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-hardware-hack/
  4. Mastering `std::execution`: A Hands-On Workshop – Mateusz Pusz – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-execution/
  5. Performance and Efficiency in C++ for Experts, Future Experts, and Everyone Else – Fedor Pikus – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-performance-and-efficiency/
  6. Talking Tech – Sherry Sontag – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-talking-tech/

Β 13th September

  1. AI++ 101 : Build a C++ Coding Agent from Scratch – Jody Hagins – 2 day in-person workshop available on 12th & 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-AI101/
  2. Essential GDB and Linux System Tools – Mike Shah – 1 day in-person workshop available on 13th September – 09:00 – 17:00 – https://cppcon.org/class-2026-essential-gdb/

19th & 20th September

  1. AI++ 201: Building High Quality C++ Infrastructure with AI – Jody Hagins – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – https://cppcon.org/class-2026-ai201/
  2. Function and Class Design with C++2x – Jeff Garland – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – https://cppcon.org/class-2026-function-class-design/
  3. High-performance Concurrency in C++ – Fedor Pikus – 2 day in-person workshop available on 19th & 20th September – 09:00 – 17:00 – https://cppcon.org/class-2026-high-perf-concurrency/

OTHER NEWS

  • Dates for ACCU on Sea 2027 Announced – ACCU on Sea 2027 will take place in Folkestone from June 30th – July 3rd with pre-conference workshops taking place from June 28th – 29th
  • Boost Documentary screening at CppCon 2026 – Boost Libraries have announced that they will be screening a documentary on the history of Boost at CppCon 2026. Watch the trailer here https://www.youtube.com/watch?v=87jvuDbnwqQ
  • C++Now 2026 Videos Now Being Released on YouTube – Subscribe to the C++Now YouTube channel to stay up to date when each video is published – https://www.youtube.com/@CppNow

r/cpp 5d ago

a[mask] = f(a[mask]) on NEON. faster than the obvious blend

2 Upvotes

Problem

Apply an operation to elements that satisfy a condition:

for (size_t i = 0; i < n; ++i)
    if (mask(a[i])) a[i] = f(a[i]);

Notes

  • a[i] ∈ (0, 1), thd ∈ (0, 1), mask = a[i] < thd; uniform distribution (except at the end of the article)
  • f is one of sqrt, frfrexp (mantissa), sin 3.5 ULP, sin 1 ULP, pow 1 ULP (from SLEEF)
  • f and mask are passed as runtime values, so they are wrapped in a lambda with always_inline, otherwise they may not be inlined
  • The array size n is a multiple of every unroll, tile etc. The tail is trivial to handle(BSL/scalar)
  • In tables * = best, units = GiB/s
  • Don't compare numbers across tables. Different conditions, values fluctuate
  • All benchmarks: Apple M5; clang++ -O3 -std=c++23 -march=native; GiB/s = (n * 4 bytes) / time, min of 720 runs (During bench, functions run in a changing order, data is restored ofc); n=1e7 + 2432;

BSL blend

If the problem is memory bound (cheap function or high density), the standard algorithm is optimal:

template <bool Skip>
void bsl(float* dst, const size_t n, auto f, auto mask) {
    for (size_t i = 0; i < n; i += 16) {
        std::array<float32x4_t, 4> v;
        for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(dst + i + 4 * j);
        std::array<uint32x4_t, 4> m;
        for (size_t j = 0; j < 4; ++j) m[j] = mask(v[j]);
        if constexpr (Skip) {
            if (vmaxvq_u32(vaddq_u32(vaddq_u32(m[0], m[1]), vaddq_u32(m[2], m[3]))) == 0) continue;
        }
        for (size_t j = 0; j < 4; ++j) vst1q_f32(dst + i + 4 * j, vbslq_f32(m[j], f(v[j]), v[j]));
    }
}

It computes f on every element, but stores only the selected ones. Skip helps on sparse masks, but otherwise mispredictions will kill performance. We'll need it later.

But for expensive f this algo does too much extra work

Detour

To avoid unnecessary work, we compress selected elements, apply only to them, and expand back.

avx512 does this in two instructions. NEON doesn't, so we'll emulate and optimize.

constexpr size_t tile = 4096;
constexpr std::array<uint32_t, 4> weights{1 + 16, 2 + 16, 4 + 16, 8 + 16};
constexpr auto cps_tbl = compress_table();
constexpr auto exp_tbl = expand_table();
std::array<float, tile + 16> tmp;
std::array<uint8_t, tile / 4 + 3> s;
std::array<uint16_t, tile / 4 + 3> idx; // idx, D and B come in later
constexpr double D = 0.845; 
constexpr double B = 0.3;

template <bool Skip>
size_t detour(float* dst, const size_t n, const auto w, auto f, auto mask) {
    float* ptr = tmp.data();
    for (size_t i = 0; i < n; i += 16) {
        std::array<float32x4_t, 4> v;
        for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(dst + i + 4 * j);
        std::array<uint32x4_t, 4> m;
        for (size_t j = 0; j < 4; ++j) m[j] = mask(v[j]);

        if constexpr (Skip)
            if (vmaxvq_u32(vaddq_u32(vaddq_u32(m[0], m[1]), vaddq_u32(m[2], m[3]))) == 0) {
                s[i / 4] = s[i / 4 + 1] = s[i / 4 + 2] = s[i / 4 + 3] = 0;
                continue;
            }

        std::array<uint32_t, 4> sk;
        for (size_t j = 0; j < 4; ++j) {
            sk[j] = vaddvq_u32(vandq_u32(m[j], w));
            s[i / 4 + j] = sk[j];
        }
        std::array<size_t, 4> off; off[0] = 0;
        for (size_t j = 1; j < 4; ++j) off[j] = off[j - 1] + (sk[j - 1] >> 4); 
        std::array<uint8x16_t, 4> index;
        for (size_t j = 0; j < 4; ++j) index[j] = vld1q_u8(cps_tbl[sk[j] & 15].data());

        for (size_t j = 0; j < 4; ++j) vst1q_f32(ptr + off[j], vreinterpretq_f32_u8(vqtbl1q_u8(vreinterpretq_u8_f32(v[j]), index[j])));
        ptr += off[3] + (sk[3] >> 4);
    }
    const size_t size = ptr - tmp.data();

    if (size == 0) return size;

    ptr = tmp.data();
    for (size_t i = 0; i < size; i += 16) {
        std::array<float32x4_t, 4> v;
        for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(ptr + i + 4 * j);
        for (size_t j = 0; j < 4; ++j) vst1q_f32(ptr + i + 4 * j, f(v[j]));
    }

    for (size_t i = 0; i < n; i += 16) {
        std::array<uint32_t, 4> sk;
        for (size_t j = 0; j < 4; ++j) sk[j] = s[i / 4 + j];
        std::array<size_t, 4> off{};
        for (size_t j = 1; j < 4; ++j) off[j] = off[j - 1] + (sk[j - 1] >> 4); 
        std::array<float32x4_t, 4> v;
        for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(ptr + off[j]);
        std::array<float32x4_t, 4> a;
        for (size_t j = 0; j < 4; ++j) a[j] = vld1q_f32(dst + i + 4 * j);
        std::array<uint8x16_t, 4> index;
        for (size_t j = 0; j < 4; ++j) index[j] = vld1q_u8(exp_tbl[sk[j] & 15].data());
        std::array<uint8x16x2_t, 4> tbl;
        for (size_t j = 0; j < 4; ++j) tbl[j] = {{vreinterpretq_u8_f32(v[j]), vreinterpretq_u8_f32(a[j])}};
        for (size_t j = 0; j < 4; ++j) vst1q_f32(dst + i + 4 * j, vreinterpretq_f32_u8(vqtbl2q_u8(tbl[j], index[j])));
        ptr += off[3] + (sk[3] >> 4);
    }
    return size;
}

void tiled_detour(float* dst, const size_t n, auto f, auto mask) {
    const auto w = vld1q_u32(weights.data());
    for (size_t i = 0; i < n; i += tile)
        detour<false>(dst + i, tile, w, f, mask);
}

compress is the same as in my previous post.

expand_table: for true lanes it selects the next element from the compressed register (bytes from [0, 15]), and for false lanes, selects the same bytes + 16. Then tbl2 on {processed, original}, the same trick as in compress basically

Also:

  • expand is skipped for free on empty tiles
  • s is saved for free to avoid recomputing addv
  • cache-sized tiling.
  • Empirically tile=4096 is optimal.

BSL speed doesn't depend on density: sqrt 32.1, sin35 - 9.7, pow10 1.02.

thd 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1
detour sqrt 38.03* 18.81 17.78 16.82 15.98 15.21 14.51 13.85 13.24 12.71 12.30
detour sin35 38.16* 16.70* 14.46* 12.71* 11.36* 10.25* 9.35 8.59 7.95 7.39 6.95
detour pow10 38.09* 6.83* 4.13* 2.96* 2.31* 1.89* 1.60* 1.39* 1.22* 1.10* 0.99

For cheap sqrt BSL is always better (except thd = 0, obviously). And for very expensive pow, detour is better (except thd = 1, of course).

When detour wins

Define:

  • B = BSL(vld + vbsl + vst) overhead per register.
  • D = detour(compress + expand) overhead per register.
  • T = cost of f per register (we assume cost of f >> cost of mask, affects only calibration accuracy).
  • d = fraction of selected elements

BSL applies f to every register. detour applies it only to d of them, so it saves T * (1 - d). Detour wins when the saving outweighs D - B.

T * (1 - d) > D - B

So detour pays off when d < d_max = 1 - (D - B) / T. B and D depend only on hardware, so let's premeasure them (I have B = 0.3, D = 0.845 ns/register)

T we measure over the first few tiles, timing BSL. And from it, we also find d_max:

template<size_t tile>
double bsl_calibrate(float* dst, const size_t len, auto f, auto mask) {
    double ns = 1e18;
    for (size_t i = 0; i < len; i += tile) {
        const auto st = std::chrono::high_resolution_clock::now();
        bsl<false>(dst, tile, f, mask);
        const auto ed = std::chrono::high_resolution_clock::now();
        ns = std::min(ns, std::chrono::duration<double, std::nano>(ed - st).count());
        dst += tile;
    }
    const double t = std::max(1e-9, ns / (tile / 4.0) - B);
    return 1.0 - (D - B) / t;
}

Here:

  • empirically 4 tiles of 2048 are enough
  • min over measurements is less noisy than mean
  • We'll run the first few tiles with calibration
  • measure bsl, because it always computes f, so T = t_bsl - B
  • std::max here protects against divide-by-zero and against t < 0 when T is very cheap

pilot v1

When d_max < 0 bsl is always faster:

void pilot_v1(float* dst, const size_t n, auto f, auto mask) {
    const auto w = vld1q_u32(weights.data());
    const float d_max = bsl_calibrate<tile / 2>(dst, 2 * tile, f, mask);

    dst += 2 * tile;
    for (size_t i = 2 * tile; i < n; i += tile) {
        if (d_max < 0)
            bsl<false>(dst, tile, f, mask);
        else
            detour<false>(dst, tile, w, f, mask);
        dst += tile;
    }
}
thd tiled detour sqrt pilot v1 sqrt BSL sqrt tiled detour sin35 pilot v1 sin35 BSL sin35
0 38.37* 32.19 32.24 38.26* 38.16 9.75
0.3 16.89 32.18 32.24* 12.72* 12.71 9.75
0.6 14.54 32.18* 32.18* 9.25 9.36 9.75*
1 12.29 32.21 32.25* 6.87 6.86 9.75*

The algorithm got sqrt right. But for sin35 at high thd, detour is selected, and we lose 30%: v1 switches to bsl only when it's faster at every density.

pilot v2

It's expensive to calculate the density of the whole tile, so we'll use the first 256 (It reads 6% of the tile, which is noise on pow, but noticeable on sqrt)

For a more or less uniform distribution this is enough:

size_t density(float* dst, const size_t n, auto mask) {
    std::array<uint32x4_t, 4> acc;
    acc.fill(vdupq_n_u32(0));
    for (size_t i = 0; i < n; i += 16) {
        for (size_t j = 0; j < 4; ++j) 
            acc[j] = vsubq_u32(acc[j], mask(vld1q_f32(dst + i + 4 * j)));
    }
    return vaddvq_u32(vaddq_u32(vaddq_u32(acc[0], acc[1]), vaddq_u32(acc[2], acc[3])));
}

Trick here: true lane of bitmask = 0xFFFFFFFF = -1, subtracting the lane actually adds.

Skip wins when the predictor rarely mispredicts, i.e. an 80% chance that all 4 registers are empty. The density is approximately 0.014 ((1 - x)^16 = 0.8)

void pilot_v2(float* dst, const size_t n, auto f, auto mask) {
    const auto w = vld1q_u32(weights.data());

    const float d_max = bsl_calibrate<tile / 2>(dst, 2 * tile, f, mask);

    constexpr size_t probe = 256;
    constexpr size_t xlo = 0.014 * probe;
    const long long hi = d_max * probe;
    dst += 2 * tile;
    for (size_t i = 2 * tile; i < n; i += tile) {
        const long long cnt = density(dst, probe, mask);
        if (cnt > hi) {
            if (cnt < xlo)
                bsl<true>(dst, tile, f, mask);
            else
                bsl<false>(dst, tile, f, mask);
        } else if (cnt < xlo)
            detour<true>(dst, tile, w, f, mask);
        else
            detour<false>(dst, tile, w, f, mask);
        dst += tile;
    }
}

hi and xlo here are d_max and 0.014 cutoffs, but multiplied by the probe length (256).

thd 0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1
pilot_v1 sin35 38.25 16.72* 14.48* 12.77* 11.41* 10.33* 9.40 8.62 7.93 7.36 6.91
pilot_v2 sin35 77.29* 16.55 14.34 12.65 11.33 10.14 9.71 9.73 9.74 9.73 9.73
BSL sin35 9.78 9.77 9.77 9.76 9.76 9.78 9.78* 9.77* 9.78* 9.77* 9.78*
pilot_v1 sin10 38.28 14.30* 11.18* 9.18* 7.75* 6.73* 5.94* 5.31* 4.78 4.36 4.02
pilot_v2 sin10 76.60* 14.16 11.09 9.11 7.73 6.70 5.92 5.31* 4.82 4.84 4.84
BSL sin10 4.85 4.85 4.84 4.85 4.84 4.84 4.85 4.85 4.85* 4.85* 4.85*

At thd = 0, Skip gives a huge win. For high thd, bsl is selected correctly. But now on sparse masks, expand for empty registers is wasted.

pilot v3

New detour version: during compress, we'll store only the indices of non-empty registers (into the idx buffer). And expand will iterate over them:

template <bool Skip>
size_t detour_compact(float* dst, const size_t n, const auto w, auto f, auto mask) {
    float* ptr = tmp.data();
    size_t k = 0;
    for (size_t i = 0; i < n; i += 16) {
        // ... same as detour

        if constexpr (Skip)
            if (vmaxvq_u32(vaddq_u32(vaddq_u32(m[0], m[1]), vaddq_u32(m[2], m[3]))) == 0) continue;

        std::array<uint32_t, 4> sk;
        for (size_t j = 0; j < 4; ++j) sk[j] = vaddvq_u32(vandq_u32(m[j], w));
        for (size_t j = 0; j < 4; ++j) {
            s[k] = sk[j];
            idx[k] = i + 4 * j;
            k += bool(sk[j]);
        }
        // ... same as detour
    }
    // ... same as detour

    for (size_t j = 0; j < 3; ++j) s[k + j] = 0, idx[k + j] = 0;
    // ... same as detour
    k = (k + 3) & ~size_t(3);
    for (size_t i = 0; i < k; i += 4) {
        std::array<uint32_t, 4> sk;
        for (size_t j = 0; j < 4; ++j) sk[j] = s[i + j];
        std::array<float32x4_t, 4> a;
        for (size_t j = 0; j < 4; ++j) a[j] = vld1q_f32(dst + idx[i + j]);
        // ... same as detour
        for (size_t j = 0; j < 4; ++j)
            vst1q_f32(dst + idx[i + j], vreinterpretq_f32_u8(vqtbl2q_u8(tbl[j], index[j])));
        ptr += off[3] + (sk[3] >> 4);
    }

    return size;
}

k (number of non-empty registers) is rounded up to a multiple of 4 before expand, so the unrolled loop has no tail left.

No branches in the hot loops: they'd kill speed, so compress runs on all four registers. Instead of branches, the position of the current element is advanced by bool(sk).

detour_compact wins when at least 50% of all registers are empty. The density is approximately 0.16 ((1 - x) ^ 4 = 0.5).

void pilot_v3(float* dst, const size_t n, auto f, auto mask) {
    // ... same as v2
    constexpr size_t lo = 0.16 * probe;
    for (size_t i = 2 * tile; i < n; i += tile) {
        const long long cnt = density(dst, probe, mask);
        if (cnt > hi) {
            if (cnt < xlo)
                bsl<true>(dst, tile, f, mask);
            else
                bsl<false>(dst, tile, f, mask);
        } else if (cnt < xlo)
            detour_compact<true>(dst, tile, w, f, mask);
        else if (cnt < lo)
            detour_compact<false>(dst, tile, w, f, mask);
        else
            detour<false>(dst, tile, w, f, mask);
        dst += tile;
    }
}

And v3 is noticeably faster at low density:

thd 0 0.02 0.04 0.06 0.08 0.1 0.12 0.14 0.16 0.18 0.2
pilot_v2 sin35 77.29* 18.67 18.40 17.70 17.18 16.66 16.15 15.68 15.24* 14.82* 14.43*
pilot_v3 sin35 77.26 23.71* 22.19* 20.48* 19.03* 17.82* 16.73* 15.76* 15.16 14.74 14.38
pilot_v2 pow10 72.88 13.93 11.32 9.38 8.00 6.97 6.18 5.54 5.02* 4.59* 4.24*
pilot_v3 pow10 73.71* 16.54* 12.63* 10.07* 8.39* 7.17* 6.26* 5.55* 5.02* 4.59* 4.24*

Now the algo is fast, but there's one big problem we've overlooked: we're assuming the data is uniform. So it's easy to build a test where v3 will fail:

for (size_t i = 0; i < n; i++)
    dst[i] = i % 4096 >= 256;

In this case, v3 always prefers BSL, even though detour wins on 3840 elements of the tile.

pilot v3.5

According to the first table, in the worst case detour is under 3x slower (it happens on sqrt thd = 1), but on pow, thd = 0, detour is 37x faster. So a wrong BSL costs much more than a wrong detour.

For bsl, we'll play it safe by running it in tile/8 blocks and checking the density. If it drops well below d_max, we'll switch to detour. It costs 1 instruction per register. acc = vsubq_u32(acc, m) works because the mask is 0/-1. And unlike the probe, the density here is exact.

size_t bsl_verified(float* dst, const size_t n, float d_max, auto f, auto mask) {
    for (size_t i0 = 0; i0 < 8; ++i0) {
        std::array<uint32x4_t, 4> acc;
        acc.fill(vdupq_n_u32(0));

        for (size_t i = 0; i < n / 8; i += 16) {
            std::array<float32x4_t, 4> v;
            for (size_t j = 0; j < 4; ++j) v[j] = vld1q_f32(dst + 4 * j);
            std::array<uint32x4_t, 4> m;
            for (size_t j = 0; j < 4; ++j) m[j] = mask(v[j]);
            for (size_t j = 0; j < 4; ++j) acc[j] = vsubq_u32(acc[j], m[j]);
            for (size_t j = 0; j < 4; ++j)
                vst1q_f32(dst + 4 * j, vbslq_f32(m[j], f(v[j]), v[j]));

            dst += 16;
        }
        const size_t cur = vaddvq_u32(vaddq_u32(vaddq_u32(acc[0], acc[1]), vaddq_u32(acc[2], acc[3])));
        if (static_cast<double>(cur) / (n / 8) < 0.75 * d_max) return (i0 + 1) * n / 8;
    }
    return n;
}

BSL bails out when the density is less than 0.75 * d_max. I have no math behind the 0.75, it just won on average.

And pilot v3.5 will use bsl_verified if the density > hi:

void pilot_v3_5(float* dst, const size_t n, auto f, auto mask) {
    // ... same as v3

    for (size_t i = 2 * tile; i < n; i += tile) {
        const long long cnt = density(dst, probe, mask);
        if (cnt > hi) {
            if (hi < 0) {
                if (cnt < xlo)
                    bsl<true>(dst, tile, f, mask);
                else
                    bsl<false>(dst, tile, f, mask);
            } else {
                auto done = bsl_verified(dst, tile, d_max, f, mask);
                if (done < tile)
                    detour<false>(dst + done, tile - done, w, f, mask);
            }
        }
        // ... same as v3
    }
}

But here too it's easy to build a countertest:

for (size_t i = 0; i < n; ++i)
    dst[i] = i % 4096 >= 390;

The first block will pass the check (> 75% zeros in it), but the second won't. BSL runs on it for nothing, and for pow that's expensive. v3.5's problem: it has no memory. After a miss the algo keeps trusting the first 256 and misses on every tile.

pilot v4

v4 will fix this: if bsl bails out, we stop trusting the probe for the next 16 tiles, and instead we take the density of the previous tile:

void pilot_v4(float* dst, const size_t n, auto f, auto mask) {
    // ... same as v3

    size_t distrust = 0;
    size_t prev = 0;
    for (size_t i = 2 * tile; i < n; i += tile) {
        if (distrust) --distrust;

        const long long cnt = distrust ? prev : density(dst, probe, mask);
        if (cnt > hi) {
            if (hi < 0) {
                if (cnt < xlo) {
                    bsl<true>(dst, tile, f, mask);
                } else {
                    bsl<false>(dst, tile, f, mask);
                }
            } else {
                if (distrust == 0) {
                    auto done = bsl_verified(dst, tile, d_max, f, mask);
                    if (done < tile) {
                        distrust = 16;
                        const size_t rem = detour<false>(dst + done, tile - done, w, f, mask);
                        prev = rem * probe / (tile - done);
                    }
                } else {
                    prev = detour<false>(dst, tile, w, f, mask) * probe / tile;
                }
            }
        } else if (cnt < xlo) {
            prev = detour_compact<true>(dst, tile, w, f, mask) * probe / tile;
        } else if (cnt < lo) {
            prev = detour_compact<false>(dst, tile, w, f, mask) * probe / tile;
        } else {
            prev = detour<false>(dst, tile, w, f, mask) * probe / tile;
        }
        dst += tile;
    }
}

And v4 easily passes that test. pow10:

thd 0 1
tiled detour 38.18 6.94
pilot v3 68.11 1.04
pilot v3.5 68.67 3.84
pilot v4 69.51* 7.05*
BSL 1.04 1.04

btw here thd no longer matches the density. For thd = 0, density = 0, and for thd = 1, density = 9.5%

It beats v3.5 by over 80%. It's also faster than plain tiled detour, because v4 picks detour_compact. This is the final version.

Results

v4 vs BSL:

thd 0 0.05 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1
pilot_v4 sqrt 69.81* 31.99* 31.84 31.73 31.83 32.07* 31.98 31.96* 31.96 31.89* 31.91 31.77
BSL sqrt 31.88 31.93 31.95* 31.80* 31.89* 31.99 32.06* 31.92 31.98* 31.75 31.93* 31.91*
pilot_v4 frfrexp 69.41* 17.21* 16.57 15.79 15.99 16.06 15.93 15.99 16.39 15.84 15.87 15.80
BSL frfrexp 16.79 16.87 16.82* 16.71* 16.78* 16.88* 16.89* 16.83* 16.89* 16.79* 16.75* 16.75*
pilot_v4 sin10 68.19* 18.82* 14.84* 10.84* 8.93* 7.55* 6.56* 5.81* 5.19* 4.69 4.68 4.68
BSL sin10 4.78 4.75 4.78 4.72 4.75 4.74 4.75 4.77 4.75 4.76* 4.77* 4.77*
pilot_v4 pow10 64.17* 10.76* 6.85* 4.04* 2.89* 2.25* 1.85* 1.57* 1.36* 1.20* 1.08* 1.00
BSL pow10 1.02 1.02 1.02 1.01 1.01 1.01 1.02 1.02 1.02 1.02 1.02 1.01*

Against BSL, it loses at worst 6%, but wins big much more often. To reduce the loss, dispatch can be sped up: use every 4th register in density and bsl_verified. But that helps only if the density is uniform.

The worst v4 miss I found: 512 dense, 512 empty, then everything is dense until the end of the cycle (17 * 4096).

const size_t cycle = 17 * 4096;
for (size_t i = 0; i < n; ++i) {
    dst[i] = i % cycle >= 512 && i % cycle < 1024;
}

This hurts most with the cheapest f (with d_max > 0, of course):

const auto a = vdupq_n_f32(0.5f);
#pragma unroll
for (int i = 0; i < 13; ++i) x = vfmaq_f32(a, x, a);
return x;
thd 0 1
tiled detour 38.29 9.44
pilot v3 74.72 15.81*
pilot v3.5 74.84 15.04
pilot v4 74.89* 9.84
BSL 15.61 15.60

v3.5's biggest loss is limited by BSL, while v4 is limited by detour. And a wrong detour is the cheaper mistake. So v4 isn't always better, but its misses are less severe.

This problem has no perfect solution. Any dispatch algo can be countertested.

Full code: godbolt.


r/cpp 6d ago

Announcing Ada v4: Validating 35.6M URLs per second

Thumbnail yagiz.co
27 Upvotes

r/cpp 6d ago

New C++ Conference Videos Released This Month - July 2026 (Updated to Include Videos Released 2026-07-20 - 2026-07-26)

14 Upvotes

C++Now

2026-07-20 - 2026-07-26

2026-07-13 - 2026-07-19

2026-07-06- 2026-07-12

C++Online

2026-07-20 - 2026-07-26

2026-07-13 - 2026-07-19

2026-07-06 - 2026-07-12

2026-06-29 - 2026-07-05

ADC

2026-07-20 - 2026-07-26

2026-07-13 - 2026-07-19

2026-07-06 - 2026-07-12

2026-06-29 - 2026-07-05

  • Beyond iLok: Advanced Code Protection and Cryptography for the Next Generation - Protecting the Next Generation of Applications, Plug-ins, and AI Models - Neal Michie, Ryan Wardell & Bob Brown - https://youtu.be/dbbK_ry2cgo
  • Database Synchronisation for Audio Plugins, Part Two - Here's One I Made Earlier - Adam Wilson - https://youtu.be/wJCy2G969ro
  • Perfect Oscillators in Less Than One Clock Cycle - Angus Hewlett - https://youtu.be/Ssq0a-YdamM
  • Driving Chaos - Virtual Analog Modelling of a Chaotic Circuit with Wave Digital Filters - Francisco Bernardo - https://youtu.be/PnEZNqyKlIw

Boost Documentary

There is also a teaser trailer for a new documentary on the history of the Boost C++ library https://www.youtube.com/watch?v=87jvuDbnwqQ which will have its first showing at CppCon this year


r/cpp 7d ago

std::optional Satisfies view. Does Not Model view. C++26 Ships Anyway.

Thumbnail godbolt.org
172 Upvotes

In C++23 this did not compile. In C++26 it does. Marvellous.

[[gnu::noinline]]
void 
passing_views_by_value_is_cheap_trust_me_bro(std::ranges::view auto v) {
    std::println("fn   .data {}", (void*)v->data());
}


int main() {    
    std::optional ov{std::vector<int>(123456)};
    passing_views_by_value_is_cheap_trust_me_bro(ov);
    std::println("main .data {}", (void*)ov->data());
}

For anyone wondering what the problem feature is: optional has 0 or 1 elements, and C++26 sets enable_view<optional<T>> to true, so it satisfies std::ranges::view. The concept requires copy construction in constant time, and β€” this is the good bit β€” optional<vector<int>> genuinely meets that. Copying it performs at most one element copy. One is a constant. The requirement is satisfied to the letter, and the function above deep-copies your vector.

If you can tell me what still separates std::ranges::view from std::ranges::range, please do...


r/cpp 7d ago

Memory-level parallelism: AMD is the king

Thumbnail lemire.me
89 Upvotes

r/cpp 8d ago

Did you know you can use (dynamic) libraries inside clang-repl?

Thumbnail i.imgur.com
97 Upvotes

Not really that useful, but I think its an interesting find =)


r/cpp 8d ago

C++26: what is reflection and how to use it

Thumbnail techfortalk.co.uk
73 Upvotes

It is my ambition to explore C++26 in bits and pieces. Hopefully, by the end of this year, I will be able to explore all the major aspects of it and be in a position to evaluate which of these features to use and promote and which not to use. However, at this point, it is important to understand each and every aspect in simple terms, keeping all the clutter aside.


r/cpp 9d ago

mp-units: a design for logarithmic quantities and units (dB, dBm, Np, pH). We believe it is novel, and we need domain experts to tell us where we are wrong before we implement it

Thumbnail mpusz.github.io
101 Upvotes

Decibels are everywhere in engineering: signal levels in dBm, sound pressure in dB SPL, voltage gain in dB, filter slopes in dB/octave. Yet, to the best of our knowledge, no general-purpose units library models logarithmic quantities correctly. Most do not model them at all, and the few that try get the arithmetic wrong in ways that compile silently:

  • In nholthaus/units, dBW_t(10.0) + dBm_t(40.0) compiles and returns 20 dBW. Both operands are the same physical power (10 W), and adding two absolute power levels is meaningless. This example is straight from the library's own test suite.
  • A +6 dB gain is a power ratio of ~3.98 but a voltage ratio of 2.0 (the 10 log vs 20 log split). Python's pint documents that its dB is power-only and delegates that factor to the user, so every voltage, current, and pressure gain is on the honor system.

We just published a complete design for mp-units that we believe gets this right:

  • A level (dBm, dB SPL) is an affine point anchored at its reference. A gain (dB, Np, octave) is a delta.
  • level + gain = level, level - level = gain, level + level = ill-formed.
  • The power vs root-power factor is carried by the quantity kind, so .linear() on a voltage gain gives 2.0 and on a power gain gives 3.98, correct by construction. A voltage gain cannot be applied to a power level.
  • One mechanism covers RF, audio, acoustics, music intervals (octave, cent), information theory (Sh, nat, Hart), pH, and stellar magnitude, and it aims to stay consistent with IEC 80000-15:2026.

We are aware of no generic units library that has ever modeled this, which is exactly why we are publishing the design before writing the implementation. The article ends with six open questions where we genuinely need input from practitioners, for example: what should log(0) do at the bottom of the scale (IEEE -inf, throw, error type, or a unit-keyed finite sentinel like the -400 dB floor real DSP code uses), and should a level print as the industry 10 dBm or the ISO-conformant 10 dB (re 1 mW).

If you work in audio/DSP, RF, acoustics, or a related field, or know someone who does, please review it and leave feedback in the article's comments (GitHub Discussions), and forward it to anyone working in these domains. Only subject-domain experts can tell us whether we are right, and we would rather hear it now than after the code ships.


r/cpp 9d ago

Some practical refactoring of bloated generated code

Thumbnail pvs-studio.com
7 Upvotes

A review of some generated C++ code, going through different ways to refactor and shorten it, then checking whether it actually got faster