<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://juraskrlec.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://juraskrlec.github.io/" rel="alternate" type="text/html" /><updated>2026-07-03T17:01:08+00:00</updated><id>https://juraskrlec.github.io/feed.xml</id><title type="html">juraskrlec</title><subtitle>Another site about software development. Hope you&apos;re having a great day :) </subtitle><entry><title type="html">World Cup 2026 - Humans and Technology</title><link href="https://juraskrlec.github.io/sport/technology/2026/07/03/world-cup-technology.html" rel="alternate" type="text/html" title="World Cup 2026 - Humans and Technology" /><published>2026-07-03T00:00:00+00:00</published><updated>2026-07-03T00:00:00+00:00</updated><id>https://juraskrlec.github.io/sport/technology/2026/07/03/world-cup-technology</id><content type="html" xml:base="https://juraskrlec.github.io/sport/technology/2026/07/03/world-cup-technology.html"><![CDATA[<p>After yesterday’s shock of the match between Portugal and Croatia, I needed some time to set my emotions aside and think about the final minutes. VAR technology detected that Matanović had touched the ball with his head, offside was called, and our goal for 2:2 was disallowed.</p>

<p>As the pressure on referees keeps growing, technology has been introduced into more or less every sport where millimetres matter. Unfortunately, that technology and its algorithms can completely overshadow what we all actually see. Yes - if the technology says the ball was touched, there is nothing more to discuss about that particular case. What I do want to discuss is the impact of technology on people.</p>

<p>The officiating we have now is hybrid. Technology is there primarily to help the head referee with decisions he couldn’t see. But if a referee knows that VAR can step in every time, we run into automation complacency - the tendency to lower your own attention and double-checking because you trust the system to do the job for you. It is closely related to automation bias, where a person places excessive trust in the technology’s suggestions and stops critically questioning the outcome or looking for additional information.</p>

<p>The second way of officiating is to remove technology entirely. The referee would then have to be maximally focused, knowing his career depends on his calls. But here’s what comes naturally to us: if a referee makes a mistake in that kind of system, after a while we, as people, will say, “Yes, he got it wrong, that’s human” - and everything will be fine.</p>

<p>The third way is to remove the human factor completely and let the technology itself judge every illegal touch, foul, and offside. Then we can say: fine, the algorithm has done its part, it is the referee now, and we simply have to accept it.</p>

<p>But when we have the hybrid model, people will never accept the technology as it is, because they still expect the referee to take responsibility, to see the bigger picture, rather than blindly trust the technology he’s using. And that is where the confusion, anger, and frustration come from - all the feelings yesterday’s match left us with. The referee becomes a lightning rod for a decision he didn’t actually make - he absorbs the anger for something the system determined.</p>

<p>Sadly, we will never go back to purely human officiating. Cameras now show every mistake to millions of viewers in a single second, and the commercial pressure is too great for anyone to give up the technology that “corrects” the referee. The only solution I see is to commit to technology fully and accept it as the final judge.</p>

<p>And here I have to be honest: I’m not advocating full automation because it’s infallible. The technology has its own margins of error and false positives. I advocate for it because it is <strong>consistent</strong> and because it clearly assigns responsibility. The problem with the hybrid system isn’t only about who makes the mistake, but about the mismatch between our expectations and reality: we expect human judgement, and we get an algorithmic one. When the algorithm openly becomes the referee, at least we know where we stand.</p>

<p>In tennis, almost every tournament is automated (Roland Garros is not), and in my view it works. That said, I have to be fair: a tennis line is pure geometry - the ball is either in or out. Football has far more decisions that are a matter of interpretation rather than measurement (intent on handball, interference from an offside position), so the tennis solution can’t be mapped over one-to-one. But the direction is the same: fewer grey zones, and less room for the feeling of being cheated.</p>

<p>What I feel most sorry about is the team and the people, because I know how much time and effort it takes to become even an average athlete, let alone a top one like our national team players. At this level, these things aren’t fair to everyone who lives this sport.</p>]]></content><author><name></name></author><category term="sport" /><category term="technology" /><summary type="html"><![CDATA[After yesterday’s shock of the match between Portugal and Croatia, I needed some time to set my emotions aside and think about the final minutes. VAR technology detected that Matanović had touched the ball with his head, offside was called, and our goal for 2:2 was disallowed. As the pressure on referees keeps growing, technology has been introduced into more or less every sport where millimetres matter. Unfortunately, that technology and its algorithms can completely overshadow what we all actually see. Yes - if the technology says the ball was touched, there is nothing more to discuss about that particular case. What I do want to discuss is the impact of technology on people. The officiating we have now is hybrid. Technology is there primarily to help the head referee with decisions he couldn’t see. But if a referee knows that VAR can step in every time, we run into automation complacency - the tendency to lower your own attention and double-checking because you trust the system to do the job for you. It is closely related to automation bias, where a person places excessive trust in the technology’s suggestions and stops critically questioning the outcome or looking for additional information. The second way of officiating is to remove technology entirely. The referee would then have to be maximally focused, knowing his career depends on his calls. But here’s what comes naturally to us: if a referee makes a mistake in that kind of system, after a while we, as people, will say, “Yes, he got it wrong, that’s human” - and everything will be fine. The third way is to remove the human factor completely and let the technology itself judge every illegal touch, foul, and offside. Then we can say: fine, the algorithm has done its part, it is the referee now, and we simply have to accept it. But when we have the hybrid model, people will never accept the technology as it is, because they still expect the referee to take responsibility, to see the bigger picture, rather than blindly trust the technology he’s using. And that is where the confusion, anger, and frustration come from - all the feelings yesterday’s match left us with. The referee becomes a lightning rod for a decision he didn’t actually make - he absorbs the anger for something the system determined. Sadly, we will never go back to purely human officiating. Cameras now show every mistake to millions of viewers in a single second, and the commercial pressure is too great for anyone to give up the technology that “corrects” the referee. The only solution I see is to commit to technology fully and accept it as the final judge. And here I have to be honest: I’m not advocating full automation because it’s infallible. The technology has its own margins of error and false positives. I advocate for it because it is consistent and because it clearly assigns responsibility. The problem with the hybrid system isn’t only about who makes the mistake, but about the mismatch between our expectations and reality: we expect human judgement, and we get an algorithmic one. When the algorithm openly becomes the referee, at least we know where we stand. In tennis, almost every tournament is automated (Roland Garros is not), and in my view it works. That said, I have to be fair: a tennis line is pure geometry - the ball is either in or out. Football has far more decisions that are a matter of interpretation rather than measurement (intent on handball, interference from an offside position), so the tennis solution can’t be mapped over one-to-one. But the direction is the same: fewer grey zones, and less room for the feeling of being cheated. What I feel most sorry about is the team and the people, because I know how much time and effort it takes to become even an average athlete, let alone a top one like our national team players. At this level, these things aren’t fair to everyone who lives this sport.]]></summary></entry><entry><title type="html">Left-right</title><link href="https://juraskrlec.github.io/swift/concurrency/2026/04/01/left-right.html" rel="alternate" type="text/html" title="Left-right" /><published>2026-04-01T00:00:00+00:00</published><updated>2026-04-01T00:00:00+00:00</updated><id>https://juraskrlec.github.io/swift/concurrency/2026/04/01/left-right</id><content type="html" xml:base="https://juraskrlec.github.io/swift/concurrency/2026/04/01/left-right.html"><![CDATA[<p>One day, my YouTube algorithm recommended a video by <a href="https://thesquareplanet.com">Jon Gjengset</a> called <a href="https://www.youtube.com/watch?v=tND-wBBZ8RY&amp;t=2567s">The Cost of Concurrency Coordination</a>. I strongly recommend watching it if you are interested in concurrency locking mechanisms, how they actually work at the CPU core level, and the left-right concurrency control technique - which was new to me at the time. It essentially enables wait-free read operations for any data structure. I wanted to learn it and implement it in Swift, and stumbled onto some cool things along the way.</p>

<h2 id="why-this-problem-exists-at-all">Why This Problem Exists At All</h2>

<p>Modern Apple Silicon chips have multiple cores - an M3 Pro has 12, an M2 Ultra has 24. Each core has its own L1 and L2 cache. When two cores want to read the same data, that’s fine - they can both have a copy in their local cache simultaneously. When one core wants to write, it has to invalidate every other core’s cached copy, wait for acknowledgements, then do the write. This cross-core communication is expensive and doesn’t scale - the more cores you have, the more acknowledgements you need.</p>

<p>This coordination is managed by the <a href="https://en.wikipedia.org/wiki/MESI_protocol">MESI protocol</a> - the cache coherence 
protocol used by Apple Silicon (or some MESI flavour). Every cache line is in one of four states:</p>

<ul>
  <li><strong>Modified</strong> - I have the only copy, and it’s dirty (different from RAM)</li>
  <li><strong>Exclusive</strong> - I have the only copy, and it matches RAM</li>
  <li><strong>Shared</strong> - multiple cores have this line, all matching RAM</li>
  <li><strong>Invalid</strong> - my copy is stale, someone else modified it</li>
</ul>

<p>When Core 0 wants to write to a line in <code class="language-plaintext highlighter-rouge">Shared</code> state, it broadcasts an invalidate message to all other cores, waits for acknowledgements, transitions to <code class="language-plaintext highlighter-rouge">Modified</code>, then writes. The more cores you have, the more acknowledgements you need to collect before you can write. This is the fundamental scalability problem that every concurrent data structure is working around.</p>

<p>This is why concurrent access to shared mutable state is hard. It’s not just a software problem. It’s a hardware problem. Every lock, every atomic operation, every memory barrier exists because of this physical reality.</p>

<p>The cache line is the unit of transfer. On Apple Silicon it’s 128 bytes. The CPU never moves less than 128 bytes between caches. This has a critical implication: if two completely unrelated variables happen to sit within the same 128-byte region of memory, and two different cores write to them, those cores will thrash each other’s caches even though they’re touching different data. This is false sharing - and it’s one of the most common reasons parallel code doesn’t scale.</p>

<p>I knew what a lock was. I knew what a data race was. But the idea that two variables could fight over cache just by sitting next to each other in memory - that was new to me, and honestly a bit mind-blowing.</p>

<p>This diagram is what everything else in this article is about. Each core has its own private caches - L1, L2 - and they all share L3 and RAM. When two cores read the same data, no problem. When one wants to write, it has to tell every other core to throw away their cached copy. That round-trip is what makes concurrent writes expensive, and it’s the reason locks, atomics, and memory barriers exist in the first place.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>                    ┌─────────────────────────────────────────┐
                    │              Main RAM                   │
                    │            ~100+ cycles                 │
                    └─────────────────┬───────────────────────┘
                                      │
                    ┌─────────────────┴───────────────────────┐
                    │             L3 / SLC Cache              │
                    │              ~40 cycles                 │
                    └──────────────┬──────────────────────────┘
                                   │
              ┌────────────────────┴─────────────────────┐
              │                                          │
    ┌─────────┴────────┐                       ┌─────────┴────────┐
    │      Core 0      │                       │      Core 1      │
    │                  │                       │                  │
    │  ┌────────────┐  │                       │  ┌────────────┐  │
    │  │  L2 Cache  │  │                       │  │  L2 Cache  │  │
    │  │ ~12 cycles │  │                       │  │ ~12 cycles │  │
    │  └─────┬──────┘  │                       │  └─────┬──────┘  │
    │        │         │                       │        │         │
    │  ┌─────┴──────┐  │                       │  ┌─────┴──────┐  │
    │  │  L1 Cache  │  │                       │  │  L1 Cache  │  │
    │  │  ~4 cycles │  │                       │  │  ~4 cycles │  │
    │  └─────┬──────┘  │                       │  └─────┬──────┘  │
    │        │         │                       │        │         │
    │  ┌─────┴──────┐  │                       │  ┌─────┴──────┐  │
    │  │  Registers │  │                       │  │  Registers │  │
    │  │  ~0 cycles │  │                       │  │  ~0 cycles │  │
    │  └────────────┘  │                       │  └────────────┘  │
    └──────────────────┘                       └──────────────────┘
</code></pre></div></div>

<h2 id="what-locks-actually-are">What Locks Actually Are</h2>

<p>Every lock in existence is built on one hardware primitive: <a href="https://en.wikipedia.org/wiki/Compare-and-swap">Compare-And-Swap (CAS)</a>. On ARM it’s implemented as a pair of instructions - <code class="language-plaintext highlighter-rouge">ldaxr</code> (load-acquire exclusive) and <code class="language-plaintext highlighter-rouge">stlxr</code> (store-release exclusive). The hardware guarantees these are atomic - indivisible, no intermediate state visible to other cores.
CAS does this atomically:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>if *address == expected {
    *address = new
    return success
} else {
    return failure  
}
</code></pre></div></div>

<p>If two cores race to CAS the same address, exactly one wins. The loser retries. Everything else - every lock, every mutex, every semaphore you’ve ever used - is built on top of this one instruction.</p>

<p><a href="https://developer.apple.com/documentation/os/os_unfair_lock">os_unfair_lock</a> (what <a href="https://developer.apple.com/documentation/os/osallocatedunfairlock">OSAllocatedUnfairLock</a> wraps) uses CAS to try to acquire. If it succeeds - nobody else held it - the whole thing takes ~5ns and never touches the kernel. If it fails, it spins briefly then calls into the kernel to sleep the thread. “Unfair” means no FIFO queue - when released, any waiter can grab it. This is the lock you want in Swift when you need raw performance.</p>

<p><a href="https://developer.apple.com/documentation/foundation/nslock">NSLock</a> wraps <a href="https://pubs.opengroup.org/onlinepubs/7908799/xsh/pthread_mutex_lock.html">pthread_mutex_t</a>. Heavier - ~25ns uncontended, fair (FIFO ordering), more features. For the write path of left-right it doesn’t matter, because writes are rare. I used it because it’s simple and the overhead is irrelevant when you’re only writing once every few seconds.</p>

<p>Here’s the fundamental problem with locks for reads:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>acquire lock      ← all readers queue here, one at a time
read data
release lock
</code></pre></div></div>

<p>Even if nobody is writing, readers block each other acquiring and releasing the lock. On 8 cores all trying to read the same data, 7 are waiting at any moment. Read throughput doesn’t scale with core count - it’s essentially serial. You bought an 8-core machine and you’re using one.</p>

<p><a href="https://pubs.opengroup.org/onlinepubs/009696899/functions/pthread_rwlock_rdlock.html">pthread_rwlock</a> improves this by allowing concurrent readers, but readers still have to atomically increment a shared reader count on every read. That shared counter is a single cache line being hammered by every reader on every core - bouncing between L1 caches constantly. Under high concurrency, that counter becomes your bottleneck. It’s why <code class="language-plaintext highlighter-rouge">pthread_rwlock</code> can actually be slower than a plain mutex under high read contention. You solved one problem and created another.</p>

<h2 id="the-arm-memory-model">The ARM Memory Model</h2>

<p>Here’s something that surprises most people: the CPU does not execute your instructions in the order you wrote them. Both the compiler and the CPU reorder operations to keep execution units busy. In single-threaded code this is invisible because the CPU tracks dependencies. In multi-threaded code it’s a real problem.</p>

<p>Consider Thread A:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>data = 42
ready = true
</code></pre></div></div>

<p>And Thread B:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>while !ready { }
print(data)  // Is this always 42?
</code></pre></div></div>

<p>On x86 - yes. x86’s strong memory model (<a href="https://en.wikipedia.org/wiki/Memory_ordering">TSO - Total Store Order</a>) prevents store-store reordering. On ARM - no. The CPU can reorder those two stores. Thread B might see <code class="language-plaintext highlighter-rouge">ready = true</code> but still read <code class="language-plaintext highlighter-rouge">data = 0</code>. This is not a bug. It’s a documented feature of the ARM memory model, designed to allow aggressive out-of-order execution for performance. Apple Silicon is ARM. Your M-series Mac can do this to you.</p>

<p>A memory barrier tells the CPU: do not reorder loads/stores across this point. On ARM the full barrier is <code class="language-plaintext highlighter-rouge">dmb ish</code> - expensive because it drains the store buffer and waits for acknowledgements from all other cores. You don’t want this in a hot path.
The ARM architecture gives you lighter-weight alternatives:</p>

<ul>
  <li><a href="https://developer.arm.com/documentation/ddi0596/2020-12/Base-Instructions/LDAR--Load-Acquire-Register-">ldar</a> - load-acquire: this load cannot be reordered with any memory operation after it</li>
  <li><a href="https://developer.arm.com/documentation/ddi0596/2021-06/Base-Instructions/STLR--Store-Release-Register-">stlr</a> - store-release: this store cannot be reordered with any memory operation before it</li>
</ul>

<p>These are cheaper than a full barrier but sufficient for most synchronization patterns - specifically the acquire-release pairing.</p>

<h2 id="the-left-right-algorithm">The Left-Right Algorithm</h2>

<p>The insight behind left-right is simple: what if readers never had to touch any shared mutable state at all? No lock to acquire, no counter to increment, no cache line to contend on. Just data, sitting there, waiting to be read.</p>

<p>The way you achieve this is by keeping two complete copies of your data structure. At any moment, one copy is active - readers use it. The other is inactive - the writer owns it. Readers never touch the inactive copy. The writer never touches the active copy while readers are on it.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>         ┌──────────────┐         ┌──────────────┐
Readers  │    ACTIVE    │         │   INACTIVE   │  Writer
  ──────▶│    Copy A    │         │    Copy B    │◀──────
  ──────▶│              │         │              │
  ──────▶└──────────────┘         └──────────────┘
                    ▲
                    │
               readIndex
          (atomic bool - tells
          readers which copy
              to use)
</code></pre></div></div>

<h3 id="the-write-sequence">The Write Sequence</h3>

<p>Step 1. The writer applies the operation to the inactive copy (B). Readers are all on A - this is completely safe, no coordination needed.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>         ┌──────────────┐         ┌──────────────┐
Readers  │    ACTIVE    │         │   INACTIVE   │  Writer
  ──────▶│    Copy A    │         │    Copy B'   │◀── writes here
  ──────▶│  (old value) │         │  (new value) │
         └──────────────┘         └──────────────┘
</code></pre></div></div>

<p>Step 2. The writer atomically flips <code class="language-plaintext highlighter-rouge">readIndex</code>. New readers now go to B’ - the updated copy. Old readers that were already reading A continue until they finish. You can’t stop them mid-read.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>         ┌──────────────┐         ┌──────────────┐
Old      │  now         │         │  now ACTIVE  │  New readers
readers  │  INACTIVE    │         │    Copy B'   │◀──────
finishing│    Copy A    │         │  (new value) │◀──────
  ──────▶│  (old value) │         │              │
         └──────────────┘         └──────────────┘
</code></pre></div></div>

<p>Step 3. The writer waits for all readers that were on A to finish. This is called the drain. Once drained, A is truly inactive - no thread is touching it.</p>

<p>Step 4. The writer applies the same operation to A. Both copies are now identical and up to date.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>         ┌──────────────┐         ┌──────────────┐
         │   INACTIVE   │         │    ACTIVE    │  Readers
Writer ──▶│   Copy A'   │         │    Copy B'   │◀──────
         │  (new value) │         │  (new value) │◀──────
         └──────────────┘         └──────────────┘
</code></pre></div></div>

<p>The next write starts from this state, with A’ as the inactive copy. The two sides alternate on every write.</p>

<h3 id="the-drain-problem">The Drain Problem</h3>

<p>After the flip, the writer needs to wait for readers that were mid-read on the old copy to finish. But how does it know when they’re done - without shared mutable state?</p>

<p>This is the clever part. Each reader gets its own epoch counter - a single integer that lives on its own private cache line. The protocol is simple:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Even value → not currently reading
Odd value  → currently reading
</code></pre></div></div>

<p>The read protocol:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>increment my counter  (even → odd)   "I am starting a read"
load readIndex                        "which copy do I use?"
read the data
increment my counter  (odd → even)   "I am done"
</code></pre></div></div>

<p>After the flip, the writer snapshots all epoch counters. Any counter that was odd at the moment of the flip belongs to a reader that was mid-read on the old copy. The writer spins on those specific counters until they go even. Readers that start after the flip go to the new copy - the writer doesn’t care about them.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Writer after flip:

epoch[0] = 4  (even) → not reading, skip
epoch[1] = 7  (odd)  → was reading on old copy, wait...
epoch[2] = 2  (even) → not reading, skip
epoch[3] = 11 (odd)  → was reading on old copy, wait...

... epoch[1] becomes 8, epoch[3] becomes 12 ...

All clear. Apply operation to now-inactive copy.
</code></pre></div></div>

<p>Each reader writes only to its own epoch counter - its own private cache line. No other thread writes to it. The writer only reads the counters during the drain, it never writes them. So there is no cache line bouncing between cores during the read path.</p>

<p>This is the whole algorithm. Two copies, an atomic flip, epoch counters, a drain. Everything else in the implementation is making this correct and efficient in Swift.</p>

<h2 id="swift-atomics">Swift Atomics</h2>

<p>Swift has no built-in atomic operations. For a long time, if you needed atomics in Swift you had to drop into C or use deprecated OSAtomic APIs. <a href="https://github.com/apple/swift-atomics">swift-atomics</a> is Apple’s official answer to that - a package that exposes atomic operations on primitive types with the full C++ memory model, available from pure Swift.</p>

<h3 id="managedatomic-vs-unsafeatomic">ManagedAtomic vs UnsafeAtomic</h3>

<p>There are two atomic types you’ll use. The difference is ownership of storage.
<code class="language-plaintext highlighter-rouge">ManagedAtomic&lt;T&gt;</code> is a class. Heap allocated, ARC managed. When you write:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>let counter = ManagedAtomic&lt;Int&gt;(0)
</code></pre></div></div>

<p>What you get in memory:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Stack:
┌──────────────────┐
│  counter (ptr)   │──────┐  8 bytes, a pointer
└──────────────────┘      │
                          ▼
Heap:
┌──────────────────────────────────┐
│  ARC refcount    │  8 bytes      │
├──────────────────────────────────┤
│  type metadata   │  8 bytes      │
├──────────────────────────────────┤
│  atomic value    │  8 bytes      │  ← the actual Int
└──────────────────────────────────┘
</code></pre></div></div>

<p>Every operation requires following that pointer. One extra memory access on every read or write. For a single atomic value this is fine - the overhead is negligible. For an array of atomics where layout matters, it’s a problem.</p>

<p><code class="language-plaintext highlighter-rouge">UnsafeAtomic&lt;T&gt;</code> does not own its storage. You provide the memory, you manage the lifetime. The key type is <code class="language-plaintext highlighter-rouge">UnsafeAtomic&lt;T&gt;.Storage</code> - a fixed-size chunk of properly aligned bytes that you embed directly in your own struct:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>struct EpochCounter {
    var storage: UnsafeAtomic&lt;UInt&gt;.Storage = .init(0)
}
</code></pre></div></div>

<p>When you put these in an array, the atomic values are contiguous in memory exactly where you put them. This is what lets you control layout - essential for cache line padding.</p>

<p>To operate on the storage you create a temporary handle:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>withUnsafeMutablePointer(to: &amp;storage) { ptr in
    UnsafeAtomic&lt;UInt&gt;(at: ptr).wrappingIncrement(ordering: .releasing)
}
</code></pre></div></div>

<p>The handle is just a pointer wrapper. It doesn’t own anything, doesn’t allocate anything. It exists for the duration of the closure and disappears. The value lives in <code class="language-plaintext highlighter-rouge">storage</code>.</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Array of EpochCounters with UnsafeAtomic:

┌──────────────────────────────────┐
│  UInt value  │  120 bytes pad   │  ← EpochCounter[0], 128 bytes, own cache line
├──────────────────────────────────┤
│  UInt value  │  120 bytes pad   │  ← EpochCounter[1], 128 bytes, own cache line
└──────────────────────────────────┘
No pointers. No heap. Values exactly where you need them.

Array of EpochCounters with ManagedAtomic:

┌─────┬─────┬─────┬─────┐
│ ptr │ ptr │ ptr │ ptr │  ← array of pointers, nicely laid out
└──┬──┴──┬──┴──┬──┴──┬──┘
   │     │     │     │
   ▼     ▼     ▼     ▼
  heap  heap  heap  heap   ← actual values scattered, same cache line, false sharing
</code></pre></div></div>

<p>For left-right, UnsafeAtomic is the only option that makes the cache line padding meaningful.</p>

<h3 id="memory-orderings">Memory Orderings</h3>

<p>This is the part that confused me the most at first. The ordering parameter on an atomic operation doesn’t affect whether the operation is atomic - it’s always atomic. It controls how the operation interacts with surrounding memory operations - what the CPU and compiler are allowed to reorder around it.</p>

<p>There are four you need to know:</p>

<p><code class="language-plaintext highlighter-rouge">.relaxed</code></p>

<p>No ordering constraints. The operation is atomic but the CPU can reorder it freely with respect to everything around it. Use when you only care about atomicity, not about what other memory is visible.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>nextSlot.loadThenWrappingIncrement(ordering: .relaxed)
</code></pre></div></div>

<p>On ARM this compiles to a plain atomic instruction with no barrier - fastest possible.</p>

<p><code class="language-plaintext highlighter-rouge">.acquiring</code> (loads only)</p>

<p>This load cannot be reordered with any memory operation after it. Think of it as a one-way barrier - nothing below this line can float above it.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>// ARM: ldar
epochs[i].load(ordering: .acquiring)
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">.releasing</code> (stores only)</p>

<p>This store cannot be reordered with any memory operation before it. Nothing above this line can sink below it.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>// ARM: stlr
epoch.wrappingIncrement(ordering: .releasing)
</code></pre></div></div>

<p><strong>The acquire-release pairing</strong> is the fundamental synchronization mechanism in left-right. A <code class="language-plaintext highlighter-rouge">.releasing</code> store on Thread A synchronizes with an <code class="language-plaintext highlighter-rouge">.acquiring</code> load of the same variable on Thread B. Once that synchronization is established, Thread B is guaranteed to see all memory writes Thread A did before the release.</p>

<p>Without the pairing:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Thread A:
data = 42
ready.store(true, .relaxed)   ← no ordering guarantee

Thread B:
ready.load(.relaxed)          ← sees true
print(data)                   ← might still see 0 on ARM
</code></pre></div></div>

<p>With the pairing:</p>
<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>Thread A:
data = 42
ready.store(true, .releasing)  ← stlr: data write cannot sink below this

Thread B:
ready.load(.acquiring)         ← ldar: nothing below can float above this
print(data)                    ← guaranteed to see 42
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">stlr</code>/<code class="language-plaintext highlighter-rouge">ldar</code> pair establishes a happens-before relationship across threads. This is exactly the same hardware mechanism described in the ARM memory model section - now you’re using it directly from Swift.</p>

<h3 id="how-this-maps-to-left-right">How This Maps to Left-Right</h3>

<p>Now every ordering choice in the implementation has a concrete reason:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>// Reader entering - stlr
// prevents the data read below from floating above this signal
cell.epochs[slot].increment()  // .releasing

// Which copy? - plain ldr
// the .releasing increment above already acts as a barrier
let useRight = cell.readIndex.load(ordering: .relaxed)

// Read the data
let result = body(useRight ? cell.right : cell.left)

// Reader done - stlr
// prevents the data read above from sinking below this signal
cell.epochs[slot].increment()  // .releasing
</code></pre></div></div>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>// Writer flipping readIndex - stlr
// the inactive copy write cannot become visible after this flip
readIndex.store(inactiveIndex, ordering: .releasing)

// Writer polling during drain - ldar
// synchronizes with reader's .releasing increment
// guarantees writer sees reader's data access as complete
while epochs[i].load(ordering: .acquiring) == seenEpoch { }
</code></pre></div></div>

<p>Without <code class="language-plaintext highlighter-rouge">.releasing</code> on the reader’s exit increment, the CPU could reorder the data read to after the “I’m done” signal - the writer would think the reader finished when it hadn’t. Without <code class="language-plaintext highlighter-rouge">.acquiring</code> on the writer’s drain poll, the CPU could speculate past the while loop before the reader’s stores were actually visible.</p>

<p>On ARM these compile to <code class="language-plaintext highlighter-rouge">stlr</code> and <code class="language-plaintext highlighter-rouge">ldar</code> - lightweight, no full <code class="language-plaintext highlighter-rouge">dmb ish</code> barrier anywhere in the hot path. On x86 they compile to plain <code class="language-plaintext highlighter-rouge">mov</code> instructions because TSO gives you acquire-release semantics for free. <code class="language-plaintext highlighter-rouge">swift-atomics</code> handles the difference transparently.</p>

<h2 id="the-swift-implementation">The Swift Implementation</h2>

<p>Now that we understand the algorithm, the hardware, and the primitives, the implementation becomes straightforward. Every line has a reason.</p>

<h3 id="epochcounter">EpochCounter</h3>

<p>The first thing we need is the epoch counter. One per reader, padded to a full cache line so there’s no false sharing between slots.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>struct EpochCounter {
    var storage: UnsafeAtomic&lt;UInt&gt;.Storage = .init(0)

    #if arch(arm64)
    private let _pad: (UInt64, UInt64, UInt64, UInt64,
                       UInt64, UInt64, UInt64, UInt64,
                       UInt64, UInt64, UInt64, UInt64,
                       UInt64, UInt64, UInt64) =
        (0,0,0,0,0,0,0,0,0,0,0,0,0,0,0)
    #else
    private let _pad: (UInt64, UInt64, UInt64,
                       UInt64, UInt64, UInt64, UInt64) =
        (0,0,0,0,0,0,0)
    #endif

    init() {}

    func increment() {
        withUnsafeMutablePointer(to: &amp;storage) { ptr in
            UnsafeAtomic&lt;UInt&gt;(at: ptr).wrappingIncrement(ordering: .releasing)
        }
    }

    func load() -&gt; UInt {
        withUnsafeMutablePointer(to: &amp;storage) { ptr in
            UnsafeAtomic&lt;UInt&gt;(at: ptr).load(ordering: .acquiring)
        }
    }
}
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">UnsafeAtomic&lt;UInt&gt;.Storage</code> is 8 bytes embedded directly in the struct - no heap, no pointer. The padding fills the rest of the cache line: 120 bytes on ARM64 (128 - 8), 56 bytes on x86 (64 - 8). The <code class="language-plaintext highlighter-rouge">#if</code> arch block handles both platforms.</p>

<p><code class="language-plaintext highlighter-rouge">increment()</code> uses <code class="language-plaintext highlighter-rouge">.releasing</code> - the data read must complete before this store becomes visible to the writer. <code class="language-plaintext highlighter-rouge">load()</code> uses .acquiring - the writer synchronizes with the reader’s release and sees the data access as truly complete. These two orderings are a pair. One without the other is wrong.</p>

<p>You can verify the layout is correct at compile time:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#if arch(arm64)
assert(MemoryLayout&lt;EpochCounter&gt;.stride == 128)
#else
assert(MemoryLayout&lt;EpochCounter&gt;.stride == 64)
#endif
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">stride</code> not <code class="language-plaintext highlighter-rouge">size</code> - stride is what determines spacing between elements in an array, which is what cache line isolation depends on.</p>

<h3 id="leftright">LeftRight</h3>

<p>The container holds the two copies, the atomic flip index, the epoch counter array, and the write lock.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>public final class LeftRight&lt;T: Sendable&gt;: @unchecked Sendable {

    internal var left: T
    internal var right: T

    internal let readIndex: ManagedAtomic&lt;Bool&gt;
    internal let epochs: UnsafeMutableBufferPointer&lt;EpochCounter&gt;

    private let writeLock = NSLock()
    private let nextSlot = ManagedAtomic&lt;Int&gt;(0)

    public init(_ initial: T, maxReaders: Int = 64) {
        self.left = initial
        self.right = initial
        self.readIndex = ManagedAtomic(false)
        self.epochs = UnsafeMutableBufferPointer&lt;EpochCounter&gt;.allocate(capacity: maxReaders)
        self.epochs.initialize(repeating: EpochCounter())
    }

    deinit {
        epochs.deinitialize()
        epochs.deallocate()
    }

    public func makeReader() -&gt; Reader&lt;T&gt; {
        let slot = nextSlot.loadThenWrappingIncrement(ordering: .relaxed) % epochs.count
        return Reader(cell: self, slot: slot)
    }
}
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">@unchecked Sendable</code> - Swift can’t prove this is safe because <code class="language-plaintext highlighter-rouge">left</code> and <code class="language-plaintext highlighter-rouge">right</code> are unprotected <code class="language-plaintext highlighter-rouge">var</code> properties. You’re telling the compiler to trust you. The epoch counters and atomic orderings are what make that claim true.</p>

<p><code class="language-plaintext highlighter-rouge">ManagedAtomic&lt;Bool&gt;</code> for <code class="language-plaintext highlighter-rouge">readIndex</code> - this is fine as a class reference. There’s only one of it, read once per read operation. The pointer dereference is not your bottleneck.</p>

<p><code class="language-plaintext highlighter-rouge">UnsafeMutableBufferPointer&lt;EpochCounter&gt;</code> for the epoch array - not <code class="language-plaintext highlighter-rouge">[EpochCounter]</code>. A Swift array would give you CoW semantics and no control over memory layout. The buffer pointer gives you direct access to contiguous memory that you manage yourself. <code class="language-plaintext highlighter-rouge">allocate</code> followed by <code class="language-plaintext highlighter-rouge">initialize</code> - you need both. <code class="language-plaintext highlighter-rouge">allocate</code> reserves the memory, <code class="language-plaintext highlighter-rouge">initialize</code> puts valid <code class="language-plaintext highlighter-rouge">EpochCounter</code> values in it. Without <code class="language-plaintext highlighter-rouge">initialize</code> you have garbage values in your epoch counters and the even/odd protocol breaks immediately.</p>

<p><code class="language-plaintext highlighter-rouge">nextSlot</code> uses <code class="language-plaintext highlighter-rouge">.relaxed</code> - handing out slot indices is just a counter. No thread is synchronizing on this value, no ordering guarantees needed.</p>

<p><code class="language-plaintext highlighter-rouge">deinit</code> calls <code class="language-plaintext highlighter-rouge">deinitialize</code> then <code class="language-plaintext highlighter-rouge">deallocate</code> - in that order, always. deinitialize runs the deinitializer on each element. deallocate frees the memory.</p>

<h3 id="reader">Reader</h3>

<p>Each reader thread gets its own <code class="language-plaintext highlighter-rouge">Reader</code> - created once, kept for the lifetime of the thread. It knows its slot and holds a reference to the cell.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>public struct Reader&lt;T: Sendable&gt; {
    internal let cell: LeftRight&lt;T&gt;
    internal let slot: Int
}

extension Reader {
    public func read&lt;R&gt;(_ body: (T) -&gt; R) -&gt; R {
        cell.epochs[slot].increment()

        let useRight = cell.readIndex.load(ordering: .relaxed)

        let result = body(useRight ? cell.right : cell.left)

        cell.epochs[slot].increment()

        return result
    }
}
</code></pre></div></div>

<p>This is the hot path. Four operations:</p>

<ol>
  <li>Increment epoch counter - signals entering read, <code class="language-plaintext highlighter-rouge">.releasing </code>ensures data access can’t float above this</li>
  <li>Load <code class="language-plaintext highlighter-rouge">readIndex</code> - <code class="language-plaintext highlighter-rouge">.relaxed</code> is correct, the increment above already established ordering</li>
  <li>Read the data - the actual work, no lock, no contention</li>
  <li>Increment epoch counter again - signals done, <code class="language-plaintext highlighter-rouge">.releasing</code> ensures data access can’t sink below this</li>
</ol>

<p>No lock. No shared counter. No cache line bouncing between cores. Each reader touches only its own private epoch slot and then the data. This is why reads scale.</p>

<h3 id="the-write-path">The Write Path</h3>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>extension LeftRight {
    public func write(_ operation: (inout T) -&gt; Void) {
        writeLock.lock()
        defer { writeLock.unlock() }

        let currentReadIndex = readIndex.load(ordering: .acquiring)
        let inactiveIndex = !currentReadIndex

        if inactiveIndex == false {
            operation(&amp;left)
        } else {
            operation(&amp;right)
        }

        readIndex.store(inactiveIndex, ordering: .releasing)

        let snapshot = (0..&lt;epochs.count).map { epochs[$0].load() }

        waitForReaders(snapshot: snapshot)

        if currentReadIndex == false {
            operation(&amp;left)
        } else {
            operation(&amp;right)
        }
    }

    private func waitForReaders(snapshot: [UInt]) {
        for (i, seenEpoch) in snapshot.enumerated() {
            guard seenEpoch &amp; 1 == 1 else { continue }
            while epochs[i].load() == seenEpoch {
                sched_yield()
            }
        }
    }
}
</code></pre></div></div>

<p>Step by step:</p>

<ol>
  <li>
    <p><strong>Lock.</strong> <code class="language-plaintext highlighter-rouge">NSLock</code> enforces single-writer. In correct usage this never contends - you’re calling <code class="language-plaintext highlighter-rouge">write</code> from one thread. The lock is there to catch misuse, not to handle real contention.</p>
  </li>
  <li>
    <p><strong>Load <code class="language-plaintext highlighter-rouge">readIndex</code> with <code class="language-plaintext highlighter-rouge">.acquiring</code>.</strong> The writer needs to see the current state of the world - specifically any stores that happened before the last flip.</p>
  </li>
  <li>
    <p><strong>Apply to inactive copy.</strong> The copy readers are not on. No reader will touch it. No synchronization needed for the write itself.</p>
  </li>
  <li>
    <p><strong>Flip <code class="language-plaintext highlighter-rouge">readIndex</code> with <code class="language-plaintext highlighter-rouge">.releasing</code>.</strong> Redirects new readers to the updated copy. <code class="language-plaintext highlighter-rouge">.releasing</code> ensures the write to the inactive copy is visible to all cores before any reader gets redirected to it. Without this, a reader could arrive at the new copy before the writer’s mutations are visible - silent corruption on ARM.</p>
  </li>
  <li>
    <p><strong>Snapshot.</strong> Immediately after the flip, read all epoch counters. Any counter that is odd at this moment was mid-read on the old copy. These are the only readers the writer needs to wait for.</p>
  </li>
  <li>
    <p><strong>Drain.</strong> For each slot that was odd at snapshot time, spin until the counter changes. Any change means the reader finished. <code class="language-plaintext highlighter-rouge">sched_yield()</code> yields the thread to the OS scheduler - if the reader isn’t currently scheduled, this lets it run rather than the writer burning CPU waiting for something that can’t happen yet.</p>
  </li>
  <li>
    <p><strong>Apply to now-inactive copy.</strong> The old active copy is fully drained. Apply the same operation. Both copies are now identical.</p>
  </li>
</ol>

<h2 id="what-i-learned">What I Learned</h2>

<p>I started this as a learning exercise. I wanted to understand how left-right works, implement it in Swift, and see what I’d learn along the way. I didn’t expect to end up this deep into ARM instruction sets, cache coherence protocols, and memory ordering models.</p>

<p>A few things that stuck with me:</p>

<p><strong>False sharing was the biggest surprise.</strong> The idea that two completely unrelated variables can fight over cache just by sitting next to each other in memory - that’s not something you learn from writing everyday Swift code. You only discover it when you have to care about where things live in memory.</p>

<p><strong>Memory orderings stopped being magic words.</strong> Before this, <code class="language-plaintext highlighter-rouge">.acquiring</code> and <code class="language-plaintext highlighter-rouge">.releasing</code> felt like incantations you copy from Stack Overflow and hope for the best. After implementing left-right, they’re just ARM instructions - <code class="language-plaintext highlighter-rouge">ldar</code> and <code class="language-plaintext highlighter-rouge">stlr</code> - with a clear, specific job. I understand what breaks if I get them wrong, because I deliberately got them wrong and watched TSan catch it.</p>

<p><strong>The algorithm itself is surprisingly small.</strong> Two copies, an atomic flip, epoch counters, a drain. That’s it. The implementation is about 135 lines of Swift.</p>

<p><strong>Swift gets you surprisingly close to the metal.</strong> <code class="language-plaintext highlighter-rouge">swift-atomics</code>, <code class="language-plaintext highlighter-rouge">UnsafeMutableBufferPointer</code>, manual memory layout - these aren’t C. They’re Swift. You can write lock-free concurrent data structures in Swift without dropping into C, and the result is correct, readable, and fast.</p>

<p>The implementation is a learning exercise and is scoped accordingly - full replacement writes only, fixed reader slots, Apple Silicon focused. There’s more to build. But as a foundation for understanding how this class of primitive works, I’m happy with it.</p>

<h2 id="sources">Sources</h2>

<ul>
  <li><a href="https://github.com/juraskrlec/swift-left-right">swift-left-right</a> - my Swift implementation</li>
  <li><a href="https://concurrencyfreaks.blogspot.com/2013/12/left-right-classical-algorithm.html">Left-Right: A Classical Algorithm</a> - Pedro Ramalhete and Andreia Correia, the original description of the primitive</li>
  <li><a href="https://github.com/jonhoo/left-right">jonhoo/left-right</a> - Jon Gjengset’s Rust implementation, the direct inspiration for this project</li>
  <li><a href="https://www.youtube.com/watch?v=tND-wBBZ8RY&amp;t=2567s">The Cost of Concurrency Coordination</a> - Jon Gjengset’s video that started all of this</li>
  <li><a href="https://github.com/apple/swift-atomics">apple/swift-atomics</a> - Apple’s official Swift atomics package</li>
  <li><a href="https://developer.arm.com/documentation/ddi0487/latest">ARM Architecture Reference Manual</a> - the definitive source on <code class="language-plaintext highlighter-rouge">ldar</code>, <code class="language-plaintext highlighter-rouge">stlr</code>, and the ARM memory model</li>
</ul>]]></content><author><name></name></author><category term="swift" /><category term="concurrency" /><summary type="html"><![CDATA[One day, my YouTube algorithm recommended a video by Jon Gjengset called The Cost of Concurrency Coordination. I strongly recommend watching it if you are interested in concurrency locking mechanisms, how they actually work at the CPU core level, and the left-right concurrency control technique - which was new to me at the time. It essentially enables wait-free read operations for any data structure. I wanted to learn it and implement it in Swift, and stumbled onto some cool things along the way.]]></summary></entry><entry><title type="html">Audio Engine/JSWaveform</title><link href="https://juraskrlec.github.io/image/2024/07/18/audio-engine-jswaveform.html" rel="alternate" type="text/html" title="Audio Engine/JSWaveform" /><published>2024-07-18T00:00:00+00:00</published><updated>2024-07-18T00:00:00+00:00</updated><id>https://juraskrlec.github.io/image/2024/07/18/audio-engine-jswaveform</id><content type="html" xml:base="https://juraskrlec.github.io/image/2024/07/18/audio-engine-jswaveform.html"><![CDATA[<p>I haven’t touched Swift and SwiftUI for a while, but watching the WWDC24 videos reignited my excitement to dive deeper into Swift and SwiftUI. I’m eager to learn more about them and integrate them into all my future projects in some form. Having worked with Objective-C++ for over a decade, I’m particularly interested in understanding the differences and learning how to tackle problems using Swift, especially with the new Swift Concurrency features.</p>

<p>The addition of C++ interoperability in Swift (a topic for another post) is fantastic for transitioning legacy Objective-C++ projects to Swift. Additionally, there’s a specific area I’ve been wanting to explore for a while: audio development on iOS. So, I decided to combine my interest in Swift with AVAudioEngine, and that’s how the JSWaveform was born. You can check out the Swift package on my <a href="https://github.com/juraskrlec/JSWaveform">Github</a>.</p>

<p>In this post, I will delve into the intricacies of an Audio Engine written in Swift, leveraging the power of Apple’s AVFoundation framework. We’ll explore how to manage audio playback, apply pitch effects, and handle asynchronous audio buffering in a robust and efficient manner.</p>

<p><strong>What is AVAudioEngine?</strong></p>

<p>AVAudioEngine is an advanced audio framework that lets you build intricate audio processing chains using a graph of audio nodes. Each node performs a specific function, such as playing an audio file, applying effects, or mixing multiple audio streams. The modularity of AVAudioEngine allows you to customize and extend audio processing according to your application’s needs.</p>

<p><strong>Key Components of AVAudioEngine</strong></p>

<p>Understanding the key components of AVAudioEngine is essential for leveraging its full potential:</p>

<ol>
  <li>
    <p>Nodes: Nodes are the fundamental units in AVAudioEngine. They perform various audio tasks and can be connected to form complex audio processing graphs.</p>

    <ul>
      <li>AVAudioInputNode: Captures audio from the microphone.</li>
      <li>AVAudioOutputNode: Outputs processed audio to the device’s speakers or headphones.</li>
      <li>AVAudioPlayerNode: Plays audio files or buffers.</li>
      <li>AVAudioUnitEffect: Applies audio effects such as reverb or delay.</li>
      <li>Engine: The AVAudioEngine class manages the audio nodes and the connections between them, ensuring synchronized processing.</li>
    </ul>
  </li>
  <li>AVAudioEngine: The core class that orchestrates the entire audio processing flow.</li>
  <li>Buffers: Buffers temporarily store audio data during processing, enabling smooth transitions and real-time manipulation.</li>
</ol>

<p><strong>JSWaveform</strong></p>

<p>JSWaveform is a Swift Package that has native interfaces consisting of audio engine and pure animatable SwiftUI components in iOS, iPadOS and visionOS.</p>

<p>JSWaveform provides native Swift and SwiftUI components. For now, it has 2 major SwiftUI views:</p>

<ul>
  <li>AudioPlayerView - renders audio player which consits of play/pause button, downsampled waveform and time pitch effect button.</li>
  <li>AudioVisualizerView - renders audio visualizer which animates AudioVisualizerShape based on audio amplitudes.</li>
</ul>

<p>All the code for manipulating audio signals follows <code class="language-plaintext highlighter-rouge">JSWaveform</code>.</p>

<p><strong>Setting Up the Audio Engine</strong></p>

<p>We define an AudioEngine actor to encapsulate the functionality:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">AudioEngine</span> <span class="p">{</span>
    
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">avAudioEngine</span> <span class="o">=</span> <span class="kt">AVAudioEngine</span><span class="p">()</span>
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">audioPlayer</span> <span class="o">=</span> <span class="kt">AVAudioPlayerNode</span><span class="p">()</span>
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">audioTimePitch</span> <span class="o">=</span> <span class="kt">AVAudioUnitTimePitch</span><span class="p">()</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">audioBuffer</span><span class="p">:</span> <span class="kt">AVAudioPCMBuffer</span><span class="p">?</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">asyncBufferStream</span><span class="p">:</span> <span class="kt">AsyncStream</span><span class="o">&lt;</span><span class="kt">AVAudioPCMBuffer</span><span class="o">&gt;</span><span class="p">?</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">continuation</span><span class="p">:</span> <span class="kt">AsyncStream</span><span class="o">&lt;</span><span class="kt">AVAudioPCMBuffer</span><span class="o">&gt;.</span><span class="kt">Continuation</span><span class="p">?</span>
    
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">logger</span> <span class="o">=</span> <span class="kt">Logger</span><span class="p">(</span><span class="nv">subsystem</span><span class="p">:</span> <span class="s">"JSWaveform.AudioEngine"</span><span class="p">,</span> <span class="nv">category</span><span class="p">:</span> <span class="s">"Engine"</span><span class="p">)</span>
    
    <span class="kd">enum</span> <span class="kt">AudioEngineError</span><span class="p">:</span> <span class="kt">Error</span> <span class="p">{</span>
        <span class="k">case</span> <span class="n">bufferRetrieveError</span>
    <span class="p">}</span>
    
    <span class="nf">init</span><span class="p">()</span> <span class="p">{</span>
        <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">attach</span><span class="p">(</span><span class="n">audioPlayer</span><span class="p">)</span>
        <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">attach</span><span class="p">(</span><span class="n">audioTimePitch</span><span class="p">)</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The setup function connects the various nodes and prepares the engine for playback:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">nonisolated</span> <span class="kd">func</span> <span class="nf">setup</span><span class="p">()</span> <span class="p">{</span>
    <span class="k">let</span> <span class="nv">output</span> <span class="o">=</span> <span class="n">avAudioEngine</span><span class="o">.</span><span class="n">outputNode</span>
    <span class="k">let</span> <span class="nv">mainMixer</span> <span class="o">=</span> <span class="n">avAudioEngine</span><span class="o">.</span><span class="n">mainMixerNode</span>

    <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">audioPlayer</span><span class="p">,</span> <span class="nv">to</span><span class="p">:</span> <span class="n">audioTimePitch</span><span class="p">,</span> <span class="nv">format</span><span class="p">:</span> <span class="kc">nil</span><span class="p">)</span>
    <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">audioTimePitch</span><span class="p">,</span> <span class="nv">to</span><span class="p">:</span> <span class="n">mainMixer</span><span class="p">,</span> <span class="nv">format</span><span class="p">:</span> <span class="kc">nil</span><span class="p">)</span>
    <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">connect</span><span class="p">(</span><span class="n">mainMixer</span><span class="p">,</span> <span class="nv">to</span><span class="p">:</span> <span class="n">output</span><span class="p">,</span> <span class="nv">format</span><span class="p">:</span> <span class="kc">nil</span><span class="p">)</span>
    <span class="n">avAudioEngine</span><span class="o">.</span><span class="nf">prepare</span><span class="p">()</span>
<span class="p">}</span>
</code></pre></div></div>

<p>We prepare an asynchronous audio buffer stream to handle audio data efficiently:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">func</span> <span class="nf">prepareBuffer</span><span class="p">()</span> <span class="p">{</span>
    <span class="n">asyncBufferStream</span> <span class="o">=</span> <span class="kt">AsyncStream</span> <span class="p">{</span> <span class="n">continuation</span> <span class="k">in</span>
         <span class="k">self</span><span class="o">.</span><span class="n">continuation</span> <span class="o">=</span> <span class="n">continuation</span>
     <span class="p">}</span>
    
    <span class="n">audioPlayer</span><span class="o">.</span><span class="nf">installTap</span><span class="p">(</span><span class="nv">onBus</span><span class="p">:</span> <span class="mi">0</span><span class="p">,</span> <span class="nv">bufferSize</span><span class="p">:</span> <span class="mi">256</span><span class="p">,</span> <span class="nv">format</span><span class="p">:</span> <span class="kc">nil</span><span class="p">)</span> <span class="p">{</span> <span class="n">buffer</span><span class="p">,</span> <span class="n">_</span> <span class="k">in</span>
        <span class="k">self</span><span class="o">.</span><span class="n">continuation</span><span class="p">?</span><span class="o">.</span><span class="nf">yield</span><span class="p">(</span><span class="n">buffer</span><span class="p">)</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="kd">func</span> <span class="nf">getBuffer</span><span class="p">()</span> <span class="o">-&gt;</span> <span class="kt">AsyncStream</span><span class="o">&lt;</span><span class="kt">AVAudioPCMBuffer</span><span class="o">&gt;</span><span class="p">?</span> <span class="p">{</span>
    <span class="k">return</span> <span class="n">asyncBufferStream</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The <code class="language-plaintext highlighter-rouge">audioPlayer.installTap(onBus:bufferSize:format:block:)</code> method is an essential part of the AudioEngine class, which taps into the audio signal coming from the audioPlayer node, allowing you to process or analyze audio data in real-time. Let’s break down the components and significance of this method call.</p>

<p>The <code class="language-plaintext highlighter-rouge">installTap</code> method in AVAudioNode allows you to create a tap on a specified bus. This tap intercepts the audio data flowing through the node, providing you with a buffer of audio samples that you can use for various purposes like visualization, analysis, or further processing.</p>

<p>Parameters of the <code class="language-plaintext highlighter-rouge">installTap</code> method:</p>

<ul>
  <li>onBus: The bus number you want to tap into. For most audio nodes, bus 0 is the default output bus.</li>
  <li>bufferSize: The number of audio samples per channel in each buffer that the tap handler receives. In this case, it is set to 256.</li>
  <li>format: The audio format of the tap. Passing nil means the format of the tap will be the format of the node’s output bus.</li>
  <li>block: A block (closure) that processes the audio buffers. This block is called with each buffer of audio data.</li>
</ul>

<p>The block (closure) provided to <code class="language-plaintext highlighter-rouge">installTap</code> is where the audio processing happens:</p>

<ul>
  <li>buffer: This is the AVAudioPCMBuffer containing the audio samples tapped from the audio node.</li>
  <li>(AVAudioTime): This parameter provides the time at which the buffer’s audio data is rendered. In this case, it’s ignored.</li>
</ul>

<p>The closure yields the buffer to the <code class="language-plaintext highlighter-rouge">AsyncStream</code> continuation. This line sends the tapped audio buffer to the AsyncStream, making it available for asynchronous processing. This approach leverages Swift’s concurrency model to handle audio data efficiently, ensuring smooth real-time performance.</p>

<p><strong>Why Buffer Size 256?</strong></p>

<p>The buffer size of 256 frames (samples per channel) is chosen for several reasons:</p>

<ul>
  <li>
    <p>Latency vs. Performance Trade-Off: A smaller buffer size means lower latency, which is critical for real-time audio applications. However, smaller buffers require the CPU to handle more frequent interrupts, increasing the processing load. A buffer size of 256 is a balanced choice, providing relatively low latency without overwhelming the CPU.</p>
  </li>
  <li>
    <p>Consistency with Audio Standards: The buffer size of 256 frames is commonly used in audio processing as it aligns well with typical audio sample rates (like 44.1kHz or 48kHz), resulting in efficient processing and compatibility with audio hardware and software standards.</p>
  </li>
  <li>
    <p>Smooth Real-Time Processing: For real-time applications like audio playback, effects, and recording, using a buffer size of 256 ensures smooth operation. It allows the audio engine to deliver consistent audio data with minimal delay, enhancing the user experience in interactive audio applications.</p>
  </li>
</ul>

<p><strong>Audio Processing</strong></p>

<p>For animations, we need to processes an audio buffer to calculate the power levels (both average and peak) for each channel in the buffer.</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">func</span> <span class="nf">process</span><span class="p">(</span><span class="nv">buffer</span><span class="p">:</span> <span class="kt">AVAudioPCMBuffer</span><span class="p">)</span> <span class="p">{</span>
    <span class="k">var</span> <span class="nv">powerLevels</span> <span class="o">=</span> <span class="p">[</span><span class="kt">PowerLevels</span><span class="p">]()</span>
    <span class="k">let</span> <span class="nv">channelCount</span> <span class="o">=</span> <span class="kt">Int</span><span class="p">(</span><span class="n">buffer</span><span class="o">.</span><span class="n">format</span><span class="o">.</span><span class="n">channelCount</span><span class="p">)</span>
    <span class="k">let</span> <span class="nv">length</span> <span class="o">=</span> <span class="nf">vDSP_Length</span><span class="p">(</span><span class="n">buffer</span><span class="o">.</span><span class="n">frameLength</span><span class="p">)</span>

    <span class="k">if</span> <span class="k">let</span> <span class="nv">floatData</span> <span class="o">=</span> <span class="n">buffer</span><span class="o">.</span><span class="n">floatChannelData</span> <span class="p">{</span>
        <span class="k">for</span> <span class="n">channel</span> <span class="k">in</span> <span class="mi">0</span><span class="o">..&lt;</span><span class="n">channelCount</span> <span class="p">{</span>
            <span class="n">powerLevels</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="nf">calculatePowers</span><span class="p">(</span><span class="nv">data</span><span class="p">:</span> <span class="n">floatData</span><span class="p">[</span><span class="n">channel</span><span class="p">],</span> <span class="nv">strideFrames</span><span class="p">:</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="nv">length</span><span class="p">:</span> <span class="n">length</span><span class="p">))</span>
        <span class="p">}</span>
    <span class="p">}</span> <span class="k">else</span> <span class="k">if</span> <span class="k">let</span> <span class="nv">int16Data</span> <span class="o">=</span> <span class="n">buffer</span><span class="o">.</span><span class="n">int16ChannelData</span> <span class="p">{</span>
        <span class="k">for</span> <span class="n">channel</span> <span class="k">in</span> <span class="mi">0</span><span class="o">..&lt;</span><span class="n">channelCount</span> <span class="p">{</span>
            <span class="k">var</span> <span class="nv">floatChannelData</span><span class="p">:</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]</span> <span class="o">=</span> <span class="kt">Array</span><span class="p">(</span><span class="nv">repeating</span><span class="p">:</span> <span class="kt">Float</span><span class="p">(</span><span class="mf">0.0</span><span class="p">),</span> <span class="nv">count</span><span class="p">:</span> <span class="kt">Int</span><span class="p">(</span><span class="n">buffer</span><span class="o">.</span><span class="n">frameLength</span><span class="p">))</span>
            <span class="nf">vDSP_vflt16</span><span class="p">(</span><span class="n">int16Data</span><span class="p">[</span><span class="n">channel</span><span class="p">],</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>
            <span class="k">var</span> <span class="nv">scalar</span> <span class="o">=</span> <span class="kt">Float</span><span class="p">(</span><span class="kt">INT16_MAX</span><span class="p">)</span>
            <span class="nf">vDSP_vsdiv</span><span class="p">(</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">scalar</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>

            <span class="n">powerLevels</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="nf">calculatePowers</span><span class="p">(</span><span class="nv">data</span><span class="p">:</span> <span class="n">floatChannelData</span><span class="p">,</span> <span class="nv">strideFrames</span><span class="p">:</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="nv">length</span><span class="p">:</span> <span class="n">length</span><span class="p">))</span>
        <span class="p">}</span>
    <span class="p">}</span> <span class="k">else</span> <span class="k">if</span> <span class="k">let</span> <span class="nv">int32Data</span> <span class="o">=</span> <span class="n">buffer</span><span class="o">.</span><span class="n">int32ChannelData</span> <span class="p">{</span>
        <span class="k">for</span> <span class="n">channel</span> <span class="k">in</span> <span class="mi">0</span><span class="o">..&lt;</span><span class="n">channelCount</span> <span class="p">{</span>
            <span class="k">var</span> <span class="nv">floatChannelData</span><span class="p">:</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]</span> <span class="o">=</span> <span class="kt">Array</span><span class="p">(</span><span class="nv">repeating</span><span class="p">:</span> <span class="kt">Float</span><span class="p">(</span><span class="mf">0.0</span><span class="p">),</span> <span class="nv">count</span><span class="p">:</span> <span class="kt">Int</span><span class="p">(</span><span class="n">buffer</span><span class="o">.</span><span class="n">frameLength</span><span class="p">))</span>
            <span class="nf">vDSP_vflt32</span><span class="p">(</span><span class="n">int32Data</span><span class="p">[</span><span class="n">channel</span><span class="p">],</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>
            <span class="k">var</span> <span class="nv">scalar</span> <span class="o">=</span> <span class="kt">Float</span><span class="p">(</span><span class="kt">INT32_MAX</span><span class="p">)</span>
            <span class="nf">vDSP_vsdiv</span><span class="p">(</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">scalar</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">floatChannelData</span><span class="p">,</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>

            <span class="n">powerLevels</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="nf">calculatePowers</span><span class="p">(</span><span class="nv">data</span><span class="p">:</span> <span class="n">floatChannelData</span><span class="p">,</span> <span class="nv">strideFrames</span><span class="p">:</span> <span class="n">buffer</span><span class="o">.</span><span class="n">stride</span><span class="p">,</span> <span class="nv">length</span><span class="p">:</span> <span class="n">length</span><span class="p">))</span>
        <span class="p">}</span>
    <span class="p">}</span>
    <span class="k">self</span><span class="o">.</span><span class="n">values</span> <span class="o">=</span> <span class="n">powerLevels</span>
<span class="p">}</span>
</code></pre></div></div>

<p>If the buffer contains floating-point data, each channel’s data is processed directly. If the buffer contains 16-bit integer data, it is converted to floating-point data. The conversion is performed using the Accelerate framework’s <code class="language-plaintext highlighter-rouge">vDSP_vflt16</code> function. The data is then normalized by dividing by <code class="language-plaintext highlighter-rouge">INT16_MAX</code>. The same calculations are done if the channel contains 32-bit integer data.</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">private</span> <span class="kd">func</span> <span class="nf">calculatePowers</span><span class="p">(</span><span class="nv">data</span><span class="p">:</span> <span class="kt">UnsafePointer</span><span class="o">&lt;</span><span class="kt">Float</span><span class="o">&gt;</span><span class="p">,</span> <span class="nv">strideFrames</span><span class="p">:</span> <span class="kt">Int</span><span class="p">,</span> <span class="nv">length</span><span class="p">:</span> <span class="n">vDSP_Length</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="kt">PowerLevels</span> <span class="p">{</span>
    <span class="k">var</span> <span class="nv">max</span><span class="p">:</span> <span class="kt">Float</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="nf">vDSP_maxv</span><span class="p">(</span><span class="n">data</span><span class="p">,</span> <span class="n">strideFrames</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">max</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">max</span> <span class="o">&lt;</span> <span class="n">kMinLevel</span> <span class="p">{</span>
        <span class="n">max</span> <span class="o">=</span> <span class="n">kMinLevel</span>
    <span class="p">}</span>

    <span class="k">var</span> <span class="nv">rms</span><span class="p">:</span> <span class="kt">Float</span> <span class="o">=</span> <span class="mf">0.0</span>
    <span class="nf">vDSP_rmsqv</span><span class="p">(</span><span class="n">data</span><span class="p">,</span> <span class="n">strideFrames</span><span class="p">,</span> <span class="o">&amp;</span><span class="n">rms</span><span class="p">,</span> <span class="n">length</span><span class="p">)</span>
    <span class="k">if</span> <span class="n">rms</span> <span class="o">&lt;</span> <span class="n">kMinLevel</span> <span class="p">{</span>
        <span class="n">rms</span> <span class="o">=</span> <span class="n">kMinLevel</span>
    <span class="p">}</span>

    <span class="k">return</span> <span class="kt">PowerLevels</span><span class="p">(</span><span class="nv">average</span><span class="p">:</span> <span class="mf">20.0</span> <span class="o">*</span> <span class="nf">log10</span><span class="p">(</span><span class="n">rms</span><span class="p">),</span> <span class="nv">peak</span><span class="p">:</span> <span class="mf">20.0</span> <span class="o">*</span> <span class="nf">log10</span><span class="p">(</span><span class="n">max</span><span class="p">))</span>
<span class="p">}</span>
</code></pre></div></div>

<p>This function calculates the Root Mean Square (RMS) and peak power levels from the audio data. The function returns these values encapsulated in a <code class="language-plaintext highlighter-rouge">PowerLevels</code> struct. Here is a detailed explanation of each component:</p>

<ol>
  <li>
    <p>RMS (Root Mean Square):</p>

    <ul>
      <li>The RMS value is a measure of the average power of the audio signal.</li>
      <li>It provides a meaningful average level of the waveform’s power.</li>
      <li>The vDSP_rmsqv function from the Accelerate framework computes the RMS value of the audio data.</li>
      <li>If the RMS value is less than kMinLevel, it is set to kMinLevel to avoid taking the logarithm of zero. kMinLevel is defined as kMinLevel: Float = 0.000_000_01 // -160 dB.</li>
    </ul>
  </li>
  <li>
    <p>Peak Level:</p>

    <ul>
      <li>The peak value represents the maximum power level of the audio signal.</li>
      <li>It indicates the highest amplitude in the audio data.</li>
      <li>The vDSP_maxv function from the Accelerate framework finds the maximum value in the audio data.</li>
    </ul>
  </li>
  <li>
    <p>Decibel Conversion:</p>

    <ul>
      <li>Audio levels are often represented in decibels (dB) because the human ear perceives sound logarithmically.</li>
      <li>The conversion from linear power values to decibels is done using the logarithm function:</li>
    </ul>

    <div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code> <span class="k">let</span> <span class="nv">average</span> <span class="o">=</span> <span class="mf">20.0</span> <span class="o">*</span> <span class="nf">log10</span><span class="p">(</span><span class="n">rms</span><span class="p">)</span>
 <span class="k">let</span> <span class="nv">peak</span> <span class="o">=</span> <span class="mf">20.0</span> <span class="o">*</span> <span class="nf">log10</span><span class="p">(</span><span class="n">max</span><span class="p">)</span>
</code></pre></div>    </div>

    <ul>
      <li>Multiplying by 20 converts the logarithm of the power ratio into decibels.</li>
    </ul>
  </li>
</ol>

<p><strong>AudioVisualizer</strong></p>

<p>Now when we have some data, we need to use it. In our Model, we process audio buffer and update our observable <code class="language-plaintext highlighter-rouge">amplitudes</code>. To have smooth animations based on screen refresh rate, we need to use <code class="language-plaintext highlighter-rouge">CADisplayLink</code>.</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">public</span> <span class="kd">func</span> <span class="nf">processAudio</span><span class="p">()</span> <span class="k">async</span> <span class="p">{</span>
    <span class="k">guard</span> <span class="k">let</span> <span class="nv">bufferStream</span> <span class="o">=</span> <span class="k">await</span> <span class="n">audioEngine</span><span class="o">.</span><span class="nf">getBuffer</span><span class="p">()</span> <span class="k">else</span> <span class="p">{</span> <span class="k">return</span> <span class="p">}</span>

    <span class="k">for</span> <span class="k">await</span> <span class="n">buffer</span> <span class="k">in</span> <span class="n">bufferStream</span> <span class="p">{</span>
        <span class="k">if</span> <span class="n">isPlaying</span> <span class="p">{</span>
            <span class="n">audioProcessing</span><span class="o">.</span><span class="nf">process</span><span class="p">(</span><span class="nv">buffer</span><span class="p">:</span> <span class="n">buffer</span><span class="p">)</span>
        <span class="p">}</span> <span class="k">else</span> <span class="p">{</span>
            <span class="n">audioProcessing</span><span class="o">.</span><span class="nf">processSilence</span><span class="p">()</span>
        <span class="p">}</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="c1">// MARK: DisplayLink update method</span>
<span class="k">override</span> <span class="kd">func</span> <span class="nf">update</span><span class="p">()</span> <span class="p">{</span>
    
    <span class="k">guard</span> <span class="k">let</span> <span class="nv">levels</span> <span class="o">=</span> <span class="n">audioLevel</span><span class="o">.</span><span class="n">levelProvider</span><span class="p">?</span><span class="o">.</span><span class="n">levels</span> <span class="k">else</span> <span class="p">{</span>
        <span class="k">return</span>
    <span class="p">}</span>
    
    <span class="n">level</span> <span class="o">=</span> <span class="kt">CGFloat</span><span class="p">(</span><span class="n">levels</span><span class="o">.</span><span class="n">level</span><span class="p">)</span>
    <span class="n">peakLevel</span> <span class="o">=</span> <span class="kt">CGFloat</span><span class="p">(</span><span class="n">levels</span><span class="o">.</span><span class="n">peakLevel</span><span class="p">)</span>
    
    <span class="k">var</span> <span class="nv">targetAmplitudes</span><span class="p">:[</span><span class="kt">Double</span><span class="p">]</span>

    <span class="c1">// Example calculation</span>
    <span class="k">let</span> <span class="nv">halfCount</span> <span class="o">=</span> <span class="n">amplitudes</span><span class="o">.</span><span class="n">count</span> <span class="o">/</span> <span class="mi">2</span>
    <span class="k">let</span> <span class="nv">newAmplitudes</span> <span class="o">=</span> <span class="p">(</span><span class="mi">0</span><span class="o">..&lt;</span><span class="n">halfCount</span><span class="p">)</span><span class="o">.</span><span class="n">map</span> <span class="p">{</span> <span class="n">index</span> <span class="k">in</span>
        <span class="k">let</span> <span class="nv">normalizedIndex</span> <span class="o">=</span> <span class="kt">Double</span><span class="p">(</span><span class="n">index</span><span class="p">)</span> <span class="o">/</span> <span class="kt">Double</span><span class="p">(</span><span class="n">halfCount</span> <span class="o">-</span> <span class="mi">1</span><span class="p">)</span>
        <span class="k">let</span> <span class="nv">scalingFactor</span> <span class="o">=</span> <span class="mf">1.0</span> <span class="o">+</span> <span class="n">normalizedIndex</span> <span class="o">*</span> <span class="mf">1.5</span>
        <span class="k">let</span> <span class="nv">amplitude</span> <span class="o">=</span> <span class="n">level</span> <span class="o">*</span> <span class="p">(</span><span class="mf">1.0</span> <span class="o">-</span> <span class="n">normalizedIndex</span><span class="p">)</span> <span class="o">+</span> <span class="n">peakLevel</span> <span class="o">*</span> <span class="n">normalizedIndex</span>
        <span class="k">return</span> <span class="nf">min</span><span class="p">(</span><span class="mf">1.0</span><span class="p">,</span> <span class="nf">max</span><span class="p">(</span><span class="mf">0.0</span><span class="p">,</span> <span class="nf">pow</span><span class="p">(</span><span class="n">amplitude</span> <span class="o">*</span> <span class="n">scalingFactor</span><span class="p">,</span> <span class="mf">1.2</span><span class="p">)))</span>
    <span class="p">}</span>
    <span class="n">targetAmplitudes</span> <span class="o">=</span> <span class="n">newAmplitudes</span> <span class="o">+</span> <span class="n">newAmplitudes</span><span class="o">.</span><span class="nf">reversed</span><span class="p">()</span>

    <span class="k">if</span> <span class="n">peakLevel</span> <span class="o">&gt;</span> <span class="mi">0</span> <span class="p">{</span>
        <span class="k">self</span><span class="o">.</span><span class="n">amplitudes</span> <span class="o">=</span> <span class="k">self</span><span class="o">.</span><span class="n">amplitudes</span><span class="o">.</span><span class="nf">enumerated</span><span class="p">()</span><span class="o">.</span><span class="n">map</span> <span class="p">{</span> <span class="n">index</span><span class="p">,</span> <span class="n">current</span> <span class="k">in</span>
            <span class="k">self</span><span class="o">.</span><span class="nf">lowPassFilter</span><span class="p">(</span><span class="nv">currentValue</span><span class="p">:</span> <span class="n">current</span><span class="p">,</span> <span class="nv">targetValue</span><span class="p">:</span> <span class="n">targetAmplitudes</span><span class="p">[</span><span class="n">index</span><span class="p">])</span>
        <span class="p">}</span>
    <span class="p">}</span>
<span class="p">}</span>

<span class="kd">private</span> <span class="kd">func</span> <span class="nf">lowPassFilter</span><span class="p">(</span><span class="nv">currentValue</span><span class="p">:</span> <span class="kt">Double</span><span class="p">,</span> <span class="nv">targetValue</span><span class="p">:</span> <span class="kt">Double</span><span class="p">,</span> <span class="nv">smoothing</span><span class="p">:</span> <span class="kt">Double</span> <span class="o">=</span> <span class="mf">0.1</span><span class="p">)</span> <span class="o">-&gt;</span> <span class="kt">Double</span> <span class="p">{</span>
    <span class="k">return</span> <span class="n">currentValue</span> <span class="o">*</span> <span class="p">(</span><span class="mf">1.0</span> <span class="o">-</span> <span class="n">smoothing</span><span class="p">)</span> <span class="o">+</span> <span class="n">targetValue</span> <span class="o">*</span> <span class="n">smoothing</span>
<span class="p">}</span>
</code></pre></div></div>

<p><code class="language-plaintext highlighter-rouge">lowPassFilter</code> function is a simple yet effective method for smoothing or filtering a signal. This particular implementation is often used in audio processing, sensor data smoothing, and other applications where you want to reduce noise or fluctuations in a signal. We use it here for smooth animations.</p>

<p><img src="https://github.com/user-attachments/assets/9b0e44ae-fbc5-4316-9be1-a2b9df3ba61e" alt="AudioVisualizer" /></p>

<p><strong>AudioPlayer</strong></p>

<p>But what when you want to play audio file, and show a waveform, and skip audio while dragging on waveform like, for example, voice message?</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">func</span> <span class="nf">loadSamples</span><span class="p">(</span><span class="n">forURL</span> <span class="nv">url</span><span class="p">:</span> <span class="kt">URL</span><span class="p">,</span> <span class="n">downsampledTo</span> <span class="nv">targetSampleCount</span><span class="p">:</span> <span class="kt">Int</span><span class="p">)</span> <span class="k">async</span> <span class="k">throws</span> <span class="o">-&gt;</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]</span> <span class="p">{</span>
    <span class="n">logger</span><span class="o">.</span><span class="nf">debug</span><span class="p">(</span><span class="s">"Loading audio samples for URL: </span><span class="se">\(</span><span class="n">url</span><span class="se">)</span><span class="s">, downsampled: </span><span class="se">\(</span><span class="n">targetSampleCount</span><span class="se">)</span><span class="s">"</span><span class="p">)</span>
    
    <span class="k">let</span> <span class="nv">file</span> <span class="o">=</span> <span class="k">try</span><span class="p">?</span> <span class="kt">AVAudioFile</span><span class="p">(</span><span class="nv">forReading</span><span class="p">:</span> <span class="n">url</span><span class="p">)</span>
    <span class="k">guard</span> <span class="k">let</span> <span class="nv">format</span> <span class="o">=</span> <span class="n">file</span><span class="p">?</span><span class="o">.</span><span class="n">processingFormat</span><span class="p">,</span> <span class="k">let</span> <span class="nv">length</span> <span class="o">=</span> <span class="n">file</span><span class="p">?</span><span class="o">.</span><span class="n">length</span> <span class="k">else</span> <span class="p">{</span> <span class="k">throw</span> <span class="kt">WafevormError</span><span class="o">.</span><span class="n">audioFileNotFound</span> <span class="p">}</span>
    
    <span class="k">let</span> <span class="nv">buffer</span> <span class="o">=</span> <span class="kt">AVAudioPCMBuffer</span><span class="p">(</span><span class="nv">pcmFormat</span><span class="p">:</span> <span class="n">format</span><span class="p">,</span> <span class="nv">frameCapacity</span><span class="p">:</span> <span class="kt">AVAudioFrameCount</span><span class="p">(</span><span class="n">length</span><span class="p">))</span>
    <span class="k">try</span><span class="p">?</span> <span class="n">file</span><span class="p">?</span><span class="o">.</span><span class="nf">read</span><span class="p">(</span><span class="nv">into</span><span class="p">:</span> <span class="n">buffer</span><span class="o">!</span><span class="p">)</span>
    
    <span class="k">guard</span> <span class="k">let</span> <span class="nv">floatChannelData</span> <span class="o">=</span> <span class="n">buffer</span><span class="p">?</span><span class="o">.</span><span class="n">floatChannelData</span> <span class="k">else</span> <span class="p">{</span> <span class="k">throw</span> <span class="kt">WafevormError</span><span class="o">.</span><span class="n">bufferRetrieveError</span> <span class="p">}</span>
    <span class="k">let</span> <span class="nv">channelData</span> <span class="o">=</span> <span class="n">floatChannelData</span><span class="o">.</span><span class="n">pointee</span>
    
    <span class="k">let</span> <span class="nv">sampleCount</span> <span class="o">=</span> <span class="kt">Int</span><span class="p">(</span><span class="n">length</span><span class="p">)</span>
    <span class="k">let</span> <span class="nv">samplesPerPixel</span> <span class="o">=</span> <span class="n">sampleCount</span> <span class="o">/</span> <span class="n">targetSampleCount</span>
    <span class="k">var</span> <span class="nv">downsampledData</span> <span class="o">=</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]()</span>
    
    <span class="k">for</span> <span class="n">i</span> <span class="k">in</span> <span class="mi">0</span><span class="o">..&lt;</span><span class="n">targetSampleCount</span> <span class="p">{</span>
        <span class="k">let</span> <span class="nv">start</span> <span class="o">=</span> <span class="n">i</span> <span class="o">*</span> <span class="n">samplesPerPixel</span>
        <span class="k">let</span> <span class="nv">end</span> <span class="o">=</span> <span class="nf">min</span><span class="p">((</span><span class="n">i</span> <span class="o">+</span> <span class="mi">1</span><span class="p">)</span> <span class="o">*</span> <span class="n">samplesPerPixel</span><span class="p">,</span> <span class="n">sampleCount</span><span class="p">)</span>
        <span class="k">let</span> <span class="nv">sampleRange</span> <span class="o">=</span> <span class="n">start</span><span class="o">..&lt;</span><span class="n">end</span>
        
        <span class="k">let</span> <span class="nv">maxSample</span> <span class="o">=</span> <span class="n">sampleRange</span><span class="o">.</span><span class="n">map</span> <span class="p">{</span> <span class="n">channelData</span><span class="p">[</span><span class="nv">$0</span><span class="p">]</span> <span class="p">}</span><span class="o">.</span><span class="nf">max</span><span class="p">()</span> <span class="p">??</span> <span class="mi">0</span>
        <span class="n">downsampledData</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="n">maxSample</span><span class="p">)</span>
    <span class="p">}</span>
    
    <span class="k">return</span> <span class="n">downsampledData</span>
<span class="p">}</span>

<span class="kd">func</span> <span class="nf">loadSamples</span><span class="p">(</span><span class="nv">downsampledTo</span><span class="p">:</span> <span class="kt">Int</span><span class="p">)</span> <span class="k">async</span> <span class="k">throws</span> <span class="o">-&gt;</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]</span> <span class="p">{</span>
    <span class="k">let</span> <span class="nv">samples</span> <span class="o">=</span> <span class="k">try</span> <span class="k">await</span> <span class="n">waveformLoader</span><span class="o">.</span><span class="nf">loadSamples</span><span class="p">(</span><span class="nv">forURL</span><span class="p">:</span> <span class="n">audioURL</span><span class="p">,</span> <span class="nv">downsampledTo</span><span class="p">:</span> <span class="n">downsampledTo</span><span class="p">)</span>
    <span class="n">normalizedSamples</span> <span class="o">=</span> <span class="nf">normalizeWaveformData</span><span class="p">(</span><span class="n">samples</span><span class="p">)</span>
    <span class="k">return</span> <span class="n">normalizedSamples</span>
<span class="p">}</span>

<span class="kd">private</span> <span class="kd">func</span> <span class="nf">normalizeWaveformData</span><span class="p">(</span><span class="n">_</span> <span class="nv">data</span><span class="p">:</span> <span class="p">[</span><span class="kt">Float</span><span class="p">])</span> <span class="o">-&gt;</span> <span class="p">[</span><span class="kt">Float</span><span class="p">]</span> <span class="p">{</span>
    <span class="k">guard</span> <span class="k">let</span> <span class="nv">maxSample</span> <span class="o">=</span> <span class="n">data</span><span class="o">.</span><span class="nf">max</span><span class="p">()</span> <span class="k">else</span> <span class="p">{</span> <span class="k">return</span> <span class="p">[]</span> <span class="p">}</span>
    <span class="k">return</span> <span class="n">data</span><span class="o">.</span><span class="n">map</span> <span class="p">{</span> <span class="nv">$0</span> <span class="o">/</span> <span class="n">maxSample</span> <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>These functions loads audio data from a file, downsamples it to a specified number of samples, normalize the data and returns the downsampled data. The process involves reading the audio file, creating a buffer to hold the data, retrieving the data, calculating the appropriate downsampling parameters, and then iterating through the data to extract and store the downsampled samples. The use of the maximum value within each downsampled segment helps preserve the peaks in the audio data. You’ll get something like this:</p>

<p><img src="https://github.com/user-attachments/assets/4c7ada40-4a9d-44dd-ae11-89a9d5b92cab" alt="AudioPlayer" /></p>

<p><strong>Conclusion</strong></p>

<p>AVAudioEngine is a versatile and powerful tool that opens up a world of possibilities for iOS developers. By mastering its components and features, you can create rich and immersive audio experiences in your applications. Whether you are developing a simple audio player or a sophisticated audio processing app, AVAudioEngine provides the tools you need to bring your audio visions to life. Happy coding!</p>]]></content><author><name></name></author><category term="image" /><summary type="html"><![CDATA[I haven’t touched Swift and SwiftUI for a while, but watching the WWDC24 videos reignited my excitement to dive deeper into Swift and SwiftUI. I’m eager to learn more about them and integrate them into all my future projects in some form. Having worked with Objective-C++ for over a decade, I’m particularly interested in understanding the differences and learning how to tackle problems using Swift, especially with the new Swift Concurrency features.]]></summary></entry><entry><title type="html">Fourier Transform</title><link href="https://juraskrlec.github.io/image/2024/04/11/fourier-transform.html" rel="alternate" type="text/html" title="Fourier Transform" /><published>2024-04-11T00:00:00+00:00</published><updated>2024-04-11T00:00:00+00:00</updated><id>https://juraskrlec.github.io/image/2024/04/11/fourier-transform</id><content type="html" xml:base="https://juraskrlec.github.io/image/2024/04/11/fourier-transform.html"><![CDATA[<p>The Fourier Transform is a powerful tool for analyzing the frequency components of digital images. Originally rooted in the analysis of signals, Fourier transforms break down complex signals into simpler sinusoidal waves. When applied to images, it transforms the spatial domain of the image (i.e., the pixel intensities) into the frequency domain, revealing the image’s underlying spatial frequency components.</p>

<p><strong>Basic Principles of Discrete FT (DFT)</strong></p>

<p>The DFT is particularly useful in image processing because it provides a way to analyze the content of an image in terms of its frequency components. The basic idea is that any image can be represented as a sum of sinusoidal functions of varying magnitudes, frequencies, and phases. This transformation from spatial to frequency domain is crucial for many applications because alterations and enhancements in the frequency domain can be inversely transformed back to the spatial domain, allowing modified images to be reconstructed with desired characteristics.</p>

<p><strong>Formula</strong></p>

<p>The DFT is essentially a sampled version of the Fourier Transform and thus captures only a subset of frequencies, sufficient to comprehensively describe the original image in the spatial domain. The set of sampled frequencies in the DFT corresponds directly to the pixel count in the spatial domain image. Consequently, both the spatial domain image and its Fourier transform representation maintain the same dimensions.
Given a discrete square image F(x,y) of size NxN the DFT is:</p>

\[F(u,v) = \sum_{x=0}^{N-1} \sum_{y=0}^{N-1} f(x,y)e^{-i2\pi(\frac{xu}{N}+\frac{yv}{N})}\]

<p><strong>DFT</strong></p>

<p>Let’s see in this example. I took a picture of the tennis balls can. Using OpenCV and DFT (OpenCV uses Fast Fourier Transform - it’s an algorithm that computes DFT in \(O(NlogN)\) instead of \(O(N^2)\) time) we get this:</p>

<p><img src="https://github.com/BlinkID/blinkid-ios/assets/26868155/3da7aaab-e094-472e-ab22-89c9a77fa0f1" alt="DFT" /></p>

<p>The image is represented in a complex form, comprising both magnitude and phase components. For visualization purposes, the magnitude component is particularly significant as it reflects the strength of each frequency component, offering a clear insight into the image’s frequency spectrum. However, because the center of the image typically exhibits significantly higher values than other areas, magnitudes are often displayed on a logarithmic scale to compress the dynamic range and enhance visibility.</p>

<p>Within the Fourier domain, the central areas correspond to low-frequency components, while the outer regions denote high-frequency components. The very center of the image signifies the DC value, which is the zero frequency component representing the overall intensity of the image.</p>

<p>Compressing the dynamic range of an image can be achieved by transforming each pixel value into its logarithm. This transformation effectively enhances the visibility of lower intensity pixels. Such a technique is particularly valuable in scenarios where the dynamic range is too extensive for conventional display screens or to be captured accurately by film. The logarithmic operation functions as a straightforward point processor, applying a logarithmic curve to map the original pixel values. While the choice between using a natural logarithm or a base 10 logarithm varies, it primarily affects the scale of the output values rather than the curve’s shape. This ensures that regardless of the logarithmic base used, the compression effect on the dynamic range remains consistent. The values are then scaled to suit an 8-bit display system. The mapping function for this operation can be described as follows:</p>

\[\log_{10}(1 + ||F(u,v)||)\]

<p><strong>Inverse DFT</strong></p>

<p>The Inverse Discrete Fourier Transform (IDFT) is an essential mathematical tool in digital signal processing and image processing, serving as the counterpart to the Discrete Fourier Transform (DFT). While the DFT converts a signal or image from the spatial domain (or time domain) to the frequency domain, the IDFT performs the reverse, transforming the data back to its original spatial domain from the frequency domain. This process is crucial for applications that modify signals in the frequency domain and require a reconstruction of the original signal or image.</p>

<p>The mathematical formula for the IDFT in a two-dimensional space, typical for images, is expressed as follows for NxN images:</p>

\[F(x,y) = \frac{1}{N^2} \sum_{u=0}^{N-1} \sum_{v=0}^{N-1} f(u,v)e^{i2\pi(\frac{xu}{N}+\frac{yv}{N})}\]

<p><strong>Filters</strong></p>

<p>Filters in image processing function exactly as their name implies—they filter out unwanted elements. These filters are generally implemented as a mask array, the same size as the original image. When this mask is placed over the original image, it selectively retains only the desired attributes.</p>

<p>As previously discussed, in an image transformed by the DFT, low frequencies are located at the center, while high frequencies are dispersed around the periphery. By designing a mask array with a central circle of zeros surrounded by ones, this filter can be applied to emphasize high frequencies in the image.</p>

<p>There are mainly three types of filter used:</p>

<ul>
  <li>Low Pass Filter (LPF)</li>
  <li>High Pass Filter (HPF)</li>
  <li>Band Pass Filter (BPF)</li>
</ul>

<p><strong>Edge Detection</strong></p>

<p>Edge detection is a crucial technique in computer vision. By identifying the edges within an image, we can utilize this information for feature extraction or pattern detection purposes.</p>

<p>Edges are typically characterized by high frequencies in an image. Therefore, after performing a DFT on an image to convert it into the frequency domain, a High-Pass Filter is applied. This filter effectively blocks all low frequencies while allowing high frequencies to pass through. Subsequently, applying an inverse DFT to this filtered image reveals distinct edge features in the original image, making it a powerful tool for analysis and processing in various applications.</p>

<p>When applying HPF on the image above, and inverse DFT, we get:</p>

<p><img src="https://github.com/BlinkID/blinkid-ios/assets/26868155/f71e1961-ebf3-49bf-870e-3924a2d8c325" alt="Inverse_DFT" /></p>

<p><strong>Blur Detection</strong></p>

<p>Detecting blur in an image using the Fourier Transform involves analyzing the frequency domain representation of the image to determine the absence or presence of high-frequency components, which are typically diminished in blurred images. Below, I provide a C++ example that uses the OpenCV library to perform this task.</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>#include &lt;opencv2/opencv.hpp&gt;
#include &lt;iostream&gt;

bool isBlurry(const cv::Mat&amp; image) {
    cv::Mat gray, floatGray, dftImage, magnitudeImage;
    
    // Convert the image to grayscale
    cv::cvtColor(image, gray, cv::COLOR_BGR2GRAY);
    
    // Convert the grayscale image to float
    gray.convertTo(floatGray, CV_32F);
    
    // Compute the DFT
    cv::dft(floatGray, dftImage, cv::DFT_COMPLEX_OUTPUT);
    
    // Compute the magnitude of the complex numbers (real and imaginary)
    std::vector&lt;cv::Mat&gt; planes;
    cv::split(dftImage, planes); // planes[0] = Re(DFT(I)), planes[1] = Im(DFT(I))
    cv::magnitude(planes[0], planes[1], magnitudeImage);
    
    // Shift the DC component to the center of the image
    int cx = magnitudeImage.cols / 2;
    int cy = magnitudeImage.rows / 2;
    cv::Mat q0(magnitudeImage, cv::Rect(0, 0, cx, cy));   // Top-Left
    cv::Mat q1(magnitudeImage, cv::Rect(cx, 0, cx, cy));  // Top-Right
    cv::Mat q2(magnitudeImage, cv::Rect(0, cy, cx, cy));  // Bottom-Left
    cv::Mat q3(magnitudeImage, cv::Rect(cx, cy, cx, cy)); // Bottom-Right

    // Swap quadrants (Top-Left with Bottom-Right, Top-Right with Bottom-Left)
    cv::Mat tmp;
    q0.copyTo(tmp);
    q3.copyTo(q0);
    tmp.copyTo(q3);
    q1.copyTo(tmp);
    q2.copyTo(q1);
    tmp.copyTo(q2);
    
    // Threshold to determine the presence of significant high frequencies
    double meanValue = cv::mean(magnitudeImage)[0];
    std::cout &lt;&lt; "Mean value of DFT magnitude: " &lt;&lt; meanValue &lt;&lt; std::endl;

    // Threshold the mean value to determine if the image is blurry
    return meanValue &lt; 0.2; // Adjust the threshold based on experimentation
}

</code></pre></div></div>

<p>Let’s see what’s happening here:</p>

<ol>
  <li><strong>Grayscale Conversion</strong>: The image is converted to grayscale because the color information is not necessary for blur detection.</li>
  <li><strong>Fourier Transform</strong>: The grayscale image is converted to the frequency domain using the DFT.</li>
  <li><strong>Magnitude Calculation</strong>: The magnitude of the DFT results is calculated to analyze the frequency components.</li>
  <li><strong>Quadrant Swapping</strong>: This step shifts the zero-frequency component to the center of the spectrum, which makes analysis more intuitive.</li>
  <li><strong>Mean Value Check</strong>: By calculating the mean of the magnitude spectrum, you can infer the presence of high-frequency components. A lower mean suggests fewer high frequencies, indicative of a blur.</li>
</ol>]]></content><author><name></name></author><category term="image" /><summary type="html"><![CDATA[The Fourier Transform is a powerful tool for analyzing the frequency components of digital images. Originally rooted in the analysis of signals, Fourier transforms break down complex signals into simpler sinusoidal waves. When applied to images, it transforms the spatial domain of the image (i.e., the pixel intensities) into the frequency domain, revealing the image’s underlying spatial frequency components.]]></summary></entry><entry><title type="html">Understanding the YUV Image Format</title><link href="https://juraskrlec.github.io/ios/image/2024/02/27/yuv.html" rel="alternate" type="text/html" title="Understanding the YUV Image Format" /><published>2024-02-27T00:00:00+00:00</published><updated>2024-02-27T00:00:00+00:00</updated><id>https://juraskrlec.github.io/ios/image/2024/02/27/yuv</id><content type="html" xml:base="https://juraskrlec.github.io/ios/image/2024/02/27/yuv.html"><![CDATA[<p>In the realm of digital imaging and video processing, the YUV color format plays a crucial role. Unlike the RGB format, which is designed for human vision, YUV is tailored for efficient video compression and broadcasting. This article delves into the intricacies of the YUV image format, with a particular focus on one of its popular subformats, NV12.</p>

<p><strong>What is YUV?</strong></p>

<p>YUV refers to a color encoding system used in video applications. It separates the luminance (brightness) and chrominance (color information) into different components. This separation is advantageous for video compression, as the human eye is more sensitive to variations in brightness than color. YUV formats are extensively used in video compression algorithms, including MPEG and H.264.</p>

<p><strong>Components of YUV</strong></p>

<ul>
  <li>Y: Represents the luminance component, which captures the brightness level of the video.</li>
  <li>U and V: Represent chrominance components, which capture the color information. U indicates how much blue-colored light is in the image, minus the luminance, while V indicates how much red-colored light there is, minus the luminance.</li>
</ul>

<p><strong>YUV Subformats</strong></p>

<p>YUV has several subformats, categorized based on how the U and V components are sampled compared to the Y component. These subformats include YUV420, YUV422, and YUV444, among others. The numbers indicate the ratio of luminance to chrominance sampling, affecting the image quality and compression ratio.</p>

<p>Check out this <a href="https://fourcc.org/yuv.php#IYUV">great site</a> for great YUV pixel formats explanations. I have it bookmarked and I open it every time when I need to do something with YUVs.</p>

<p><strong>The NV12 Format</strong></p>

<p>NV12 is a specific type of YUV 4:2:0 format where chrominance components are stored as an interleaved UV plane following the Y plane. This format is particularly efficient for hardware decoding and rendering, making it widely used in mobile devices, cameras, and for streaming video content.</p>

<p><strong>Structure of NV12</strong></p>

<p>The Y component occupies the first part of the image data, with one byte per pixel, representing the luminance.
The UV component follows the Y data, with alternating U and V values. Each U and V pair is shared by four Y pixels, forming a macro pixel.
Some advantages of NV12 formats are:</p>

<ul>
  <li>Efficiency: The NV12 format is highly efficient for processing and storage, reducing bandwidth and storage requirements without significantly compromising image quality.</li>
  <li>Compatibility: NV12 is supported by a wide range of hardware decoders and media frameworks, enhancing interoperability across different platforms and devices.</li>
</ul>

<p><strong>Visualization of NV12 Format</strong></p>

<p>The best visual representation I have seen so far is the visual representation on its <a href="https://en.wikipedia.org/wiki/YCbCr">Wiki page</a>.
<img src="https://github.com/juraskrlec/juraskrlec.github.io/assets/26868155/914f2e8d-ea16-4267-8c1a-584587621f6e" alt="" /></p>

<p><strong>iOS and YUV</strong></p>

<p>When dealing with YUVs on Apple platforms, you’ll most likely use <strong>kCVPixelFormatType_420YpCbCr8BiPlanarFullRange</strong> pixel format. It is a specific type of video pixel format used within the Core Video framework on iOS and macOS platforms. It’s essential for developers working with video capture, processing, or playback applications on Apple devices. This format is closely related to the NV12 format discussed in the context of YUV color spaces, with some specific characteristics tailored to the iOS ecosystem.</p>

<p><strong>Key Characteristics</strong></p>

<ul>
  <li>Bi-Planar: The format is bi-planar, meaning the Y (luminance) and CbCr (chrominance) components are stored in two separate planes. The Y plane contains all the luminance data, with one byte per pixel. The CbCr plane contains interleaved chrominance data, with one byte for Cb (chroma blue) and one byte for Cr (chroma red) for every two pixels.</li>
  <li>Full Range: The “Full Range” in its name indicates that the luminance component (Y) uses the full 8-bit range from 0 to 255, unlike the video range (16 to 235) typically used in video encoding. This full range provides better detail in both the shadows and highlights of an image, making it suitable for high-quality video applications.</li>
  <li>420 Chroma Subsampling: This format implements 4:2:0 chroma subsampling, where the chrominance information is sampled at half the horizontal and vertical resolution of the luminance. This approach reduces the amount of data needed to represent a color image, which is beneficial for reducing file sizes and bandwidth usage without significantly impacting perceived image quality.</li>
</ul>

<p><strong>Rotation</strong></p>

<p>I had a fun work with this format recently where I needed to rotate the YUV for 90 degrees. When dealing with pixel buffers, it is important to now where the origin point. The default pixel buffer orientation on iOS is landscape right. Origin point is on the upper left corner. But what about when you change the connection’s pixel buffer video orientation or video rotation angle? Then the origin point is different. Depending on your ML models, preview layers rotation or something else, you’ll most likely rotate the YUV to its desired orientation.</p>

<p>Rotating a YUV420 image by 90 degrees involves manipulating the Y plane and the UV plane separately due to the semi-planar format of YUV420. In YUV420, the Y plane is full resolution while the UV (chroma) plane is subsampled by a factor of 2 both horizontally and vertically. The code is written in Objective-C. Really like Objective-C more than Swift, don’t hate me :)</p>

<p><strong>Common</strong></p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>size_t originalHeight = CVPixelBufferGetHeight(pixelBuffer);
size_t originalWidth = CVPixelBufferGetWidth(pixelBuffer);
size_t rotatedWidth = originalHeight; // Swap the dimensions for rotation
size_t rotatedHeight = originalWidth;
size_t originalYLumaStride = CVPixelBufferGetBytesPerRowOfPlane(pixelBuffer, 0); // For Y plane
size_t originalUVStride = CVPixelBufferGetBytesPerRowOfPlane(pixelBuffer, 1); // For the CbCr plane
        
rotatedPixelBuffer = nil;
CVReturn status = CVPixelBufferCreate(kCFAllocatorDefault, rotatedWidth, rotatedHeight, kCVPixelFormatType_420YpCbCr8BiPlanarFullRange, NULL, &amp;rotatedPixelBuffer);
(void)status;
CVPixelBufferLockBaseAddress(rotatedPixelBuffer, 0);        
</code></pre></div></div>

<p><strong>Y Plane Rotation</strong></p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>+ (void)rotateYPlaneFor:(CVImageBufferRef)pixelBuffer rotatedPixelBuffer:(CVImageBufferRef)rotatedPixelBuffer {
	uint8_t *originalYBaseAddress = (uint8_t *)CVPixelBufferGetBaseAddressOfPlane(pixelBuffer, 0);
    uint8_t *rotatedYBaseAddress = (uint8_t *)CVPixelBufferGetBaseAddressOfPlane(rotatedPixelBuffer, 0);
    for (size_t x = 0; x &lt; originalWidth; x++) {
        for (size_t y = 0; y &lt; originalHeight; y++) {
            size_t originalIndex = y * originalYLumaStride + x;
            size_t rotatedIndex;
            if (isCounterClockwise) {
                rotatedIndex = (originalWidth - x - 1) * originalHeight + y; // Counterclockwise
            }
            else {
                rotatedIndex = x * originalHeight + (originalHeight - y - 1); // Clockwise
            }
            rotatedYBaseAddress[rotatedIndex] = originalYBaseAddress[originalIndex];
        }
    }
}
</code></pre></div></div>

<p><strong>UV Plane Rotation</strong></p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>+ (void)rotateCbCrPlaneFor:(CVImageBufferRef)pixelBuffer rotatedPixelBuffer:(CVImageBufferRef)rotatedPixelBuffer {
    uint8_t *originalCbCrBaseAddress = (uint8_t *)CVPixelBufferGetBaseAddressOfPlane(pixelBuffer, 1);
    uint8_t *rotatedCbCrBaseAddress = (uint8_t *)CVPixelBufferGetBaseAddressOfPlane(rotatedPixelBuffer, 1);
    
    for (size_t x = 0; x &lt; originalWidth; x += 2) { // Half resolution for CbCr
        for (size_t y = 0; y &lt; originalHeight / 2; y++) {
            size_t originalIndex = y * originalUVStride + x;
            size_t rotatedIndex;
            if (isCounterClockwise) {
                rotatedIndex = ((originalWidth - x - 2) / 2) * originalHeight + (y * 2); // Counterclockwise
            }
            else {
                rotatedIndex =  (x / 2) * originalHeight + (originalHeight / 2 - y - 1) * 2; // Clockwise
            }

            rotatedCbCrBaseAddress[rotatedIndex] = originalCbCrBaseAddress[originalIndex];
            rotatedCbCrBaseAddress[rotatedIndex + 1] = originalCbCrBaseAddress[originalIndex + 1];
        }
    }
}
</code></pre></div></div>

<p><strong>Don’t forget</strong></p>

<p>Lock the adresses when you start:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>CVPixelBufferLockBaseAddress(pixelBuffer, 0);
CVPixelBufferLockBaseAddress(rotatedPixelBuffer, 0);
</code></pre></div></div>

<p>and release when you are done:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>CVPixelBufferUnlockBaseAddress(pixelBuffer, 0);
CFRelease(pixelBuffer);
CVPixelBufferUnlockBaseAddress(rotatedPixelBuffer, 0);
CFRelease(rotatedPixelBuffer);
</code></pre></div></div>]]></content><author><name></name></author><category term="ios" /><category term="image" /><summary type="html"><![CDATA[In the realm of digital imaging and video processing, the YUV color format plays a crucial role. Unlike the RGB format, which is designed for human vision, YUV is tailored for efficient video compression and broadcasting. This article delves into the intricacies of the YUV image format, with a particular focus on one of its popular subformats, NV12.]]></summary></entry><entry><title type="html">CMake + Swift</title><link href="https://juraskrlec.github.io/ios/swift/cmake/2023/02/20/swift-cmake.html" rel="alternate" type="text/html" title="CMake + Swift" /><published>2023-02-20T00:00:00+00:00</published><updated>2023-02-20T00:00:00+00:00</updated><id>https://juraskrlec.github.io/ios/swift/cmake/2023/02/20/swift-cmake</id><content type="html" xml:base="https://juraskrlec.github.io/ios/swift/cmake/2023/02/20/swift-cmake.html"><![CDATA[<p>It’s been a long time since I’ve posted something here. I thought that I could do a post every two weeks, but I needed a little break. And a little break turned into a bigger break. Anyway, I’ll do a post every month.</p>

<p>I had an interesting problem last year that I wanted to tackle; how to add Swift files to a framework written in Objective-C through CMake and make it work.</p>

<p>The basic concept is:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>1. Add a Swift file to your sources. This is important, if you don't Swift Compiler will be turned off. You can create an empty Swift file (dummy).
2. We need to turn on and force Swift Compiler (swiftc) for CMake to recognize its 	builds settings.
3. Add CMake SET for swiftc build settings
4. Remove other Swift flags
5. Use Swift :)
</code></pre></div></div>

<p>Let’s tackle one by one.</p>

<ul>
  <li>Add a Swift file to your sources; add this where you list your source files for cmake:</li>
</ul>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>list( APPEND &lt;your_source_list&gt;
...
&lt;path&gt;/Empty.swift
...
)
</code></pre></div></div>

<ul>
  <li>Force load <code class="language-plaintext highlighter-rouge">swiftc</code> in your <code class="language-plaintext highlighter-rouge">&lt;project&gt;.cmake</code>:</li>
</ul>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>execute_process(
    COMMAND xcrun --find swiftc
    OUTPUT_VARIABLE SWIFTC_EXECUTABLE
    OUTPUT_STRIP_TRAILING_WHITESPACE)
set( SWIFTC_FOUND TRUE )
message( STATUS "Found swiftc: ${SWIFTC_EXECUTABLE}" )

if ( SWIFTC_FOUND )
    message(STATUS "Setting swiftc: ${SWIFTC_EXECUTABLE}")
    # We need to tell cmake where swiftc is
    set( CMAKE_Swift_COMPILER ${SWIFTC_EXECUTABLE} )
    # This line is important, we need to force load it
    set( CMAKE_Swift_COMPILER_FORCED ON ) 
endif()

if( CMAKE_Swift_COMPILER )
  # Enable Swift, tell cmake there's now new swiftc build settings, without this, cmake can't properly configure attributes
  enable_language(Swift)
else()
  message(FATAL_ERROR "FATAL ERROR! Did not find swiftc?!")
endif()
</code></pre></div></div>

<ul>
  <li>Add <code class="language-plaintext highlighter-rouge">swiftc</code> build settings to your <code class="language-plaintext highlighter-rouge">&lt;project&gt;.cmake</code>:</li>
</ul>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code># Swift Features
message( STATUS "Setting Swift Settings:" )
# Swift requires modules
set( CMAKE_XCODE_ATTRIBUTE_DEFINES_MODULE "YES" )
# Enable modules for ObjectiveC - Swift
set( CMAKE_XCODE_ATTRIBUTE_CLANG_ENABLE_MODULES "YES" )
# Set Swift to ObjC Interface name
set( CMAKE_XCODE_ATTRIBUTE_SWIFT_INSTALL_OBJC_HEADER "YES")
set( CMAKE_XCODE_ATTRIBUTE_SWIFT_OBJC_INTERFACE_HEADER_NAME "${FRAMEWORK_OUTPUT_NAME}-Swift.h" )
# Set name for the source code module constructed for this target, and which will be used to import the module in implementation source files
set( CMAKE_XCODE_ATTRIBUTE_PRODUCT_MODULE_NAME "${FRAMEWORK_OUTPUT_NAME}" )
# Set paths to be searched by the Swift compiler for additional Swift modules. This is used to use ObjectiveC in Swift. Set all ObjC private headers to use for internal use in module.modulemap
set( CMAKE_XCODE_ATTRIBUTE_SWIFT_INCLUDE_PATHS "&lt;path&gt;/ProjectModule" )
# Set Swift Version
set( CMAKE_XCODE_ATTRIBUTE_SWIFT_VERSION "5.0" )
set( CMAKE_XCODE_ATTRIBUTE_SWIFT_OPTIMIZATION_LEVEL "$&lt;$&lt;CONFIG:Release&gt;:-O&gt;$&lt;$&lt;CONFIG:Debug&gt;:-Onone&gt;" )
# Set for every target
set ( CMAKE_XCODE_ATTRIBUTE_OTHER_SWIFT_FLAGS "")
</code></pre></div></div>

<ul>
  <li>Remove other Swift flags:</li>
</ul>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>set_target_properties( ${PROJECT_NAME} PROPERTIES
FRAMEWORK TRUE
...
XCODE_ATTRIBUTE_OTHER_SWIFT_FLAGS ""
XCODE_ATTRIBUTE_BUILD_LIBRARY_FOR_DISTRIBUTION "YES"
...
)
</code></pre></div></div>

<ul>
  <li>Use Swift. But be aware, you’ll need to wrap some things when using Swift in the Objective-C file if you want them private :)</li>
</ul>

<p>Feel free to message me on Twitter if you want to talk about this, or anything tech related :)</p>]]></content><author><name></name></author><category term="ios" /><category term="swift" /><category term="cmake" /><summary type="html"><![CDATA[It’s been a long time since I’ve posted something here. I thought that I could do a post every two weeks, but I needed a little break. And a little break turned into a bigger break. Anyway, I’ll do a post every month.]]></summary></entry><entry><title type="html">Swift Actors: Isolation</title><link href="https://juraskrlec.github.io/ios/swift/concurrency/2022/05/24/swift-actors-isolation.html" rel="alternate" type="text/html" title="Swift Actors: Isolation" /><published>2022-05-24T00:00:00+00:00</published><updated>2022-05-24T00:00:00+00:00</updated><id>https://juraskrlec.github.io/ios/swift/concurrency/2022/05/24/swift-actors-isolation</id><content type="html" xml:base="https://juraskrlec.github.io/ios/swift/concurrency/2022/05/24/swift-actors-isolation.html"><![CDATA[<p>By default, each mutable property and method is isolated, which means they can’t be accessed from the external code. If you want to call them, you need to use await keyword.
Immutable, constant properties and methods that you mark explicitly with nonisolated keywords can be accessed directly from the external code because they are in a non-isolated state.
Here is an example:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">GameScore</span> <span class="p">{</span>
    <span class="k">let</span> <span class="nv">name</span><span class="p">:</span> <span class="kt">String</span>
    
    <span class="c1">//...</span>
<span class="p">}</span>
</code></pre></div></div>

<p>The property <code class="language-plaintext highlighter-rouge">name</code> is immutable, so there’s no need to mark it directly with nonisolated keyword; it is safe to access from non-isolated state. Let’s add one more immutable property <code class="language-plaintext highlighter-rouge">place</code> and computed property <code class="language-plaintext highlighter-rouge">venue</code>:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>actor GameScore {
    let name: String
    let place: String
    
    var venue: String {
        "See a game: \(name) at place: \(place)"
    }
    
    //...
}
</code></pre></div></div>

<p>If we accessed the venue, we would get the following error:</p>
<blockquote>
  <p>Actor-isolated property <code class="language-plaintext highlighter-rouge">venue</code> can not be referenced from a non-isolated context</p>
</blockquote>

<p>To fix this, we must tell the compiler it is safe to access the property <code class="language-plaintext highlighter-rouge">venue</code> by marking the property nonisolated:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">GameScore</span> <span class="p">{</span>
    <span class="k">let</span> <span class="nv">name</span><span class="p">:</span> <span class="kt">String</span>
    <span class="k">let</span> <span class="nv">place</span><span class="p">:</span> <span class="kt">String</span>
    
    <span class="kd">nonisolated</span> <span class="k">var</span> <span class="nv">venue</span><span class="p">:</span> <span class="kt">String</span> <span class="p">{</span>
        <span class="s">"See a game: </span><span class="se">\(</span><span class="n">name</span><span class="se">)</span><span class="s"> at place: </span><span class="se">\(</span><span class="n">place</span><span class="se">)</span><span class="s">"</span>
    <span class="p">}</span>
    
    <span class="c1">//...</span>
<span class="p">}</span>
</code></pre></div></div>

<p>Let’s try to remove <code class="language-plaintext highlighter-rouge">venue</code> and add CustomStringConvertible protocol conformance. We will get this:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">extension</span> <span class="kt">GameScore</span><span class="p">:</span> <span class="kt">CustomStringConvertible</span> <span class="p">{</span>
    <span class="k">var</span> <span class="nv">description</span><span class="p">:</span> <span class="kt">String</span> <span class="p">{</span>
        <span class="s">"See a game: </span><span class="se">\(</span><span class="n">name</span><span class="se">)</span><span class="s"> at place: </span><span class="se">\(</span><span class="n">place</span><span class="se">)</span><span class="s">"</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>And when we try to use it, we will get the error:</p>
<blockquote>
  <p>Actor-isolated property ‘description’ cannot be used to satisfy a protocol requirement</p>
</blockquote>

<p>To fix this and to use <code class="language-plaintext highlighter-rouge">description,</code> we need to mark it with a nonisolated keyword:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">extension</span> <span class="kt">GameScore</span><span class="p">:</span> <span class="kt">CustomStringConvertible</span> <span class="p">{</span>
    <span class="kd">nonisolated</span> <span class="k">var</span> <span class="nv">description</span><span class="p">:</span> <span class="kt">String</span> <span class="p">{</span>
        <span class="s">"See a game: </span><span class="se">\(</span><span class="n">name</span><span class="se">)</span><span class="s"> at place: </span><span class="se">\(</span><span class="n">place</span><span class="se">)</span><span class="s">"</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<h2 id="conclusion">Conclusion</h2>

<p>Actors are straightforward and neat ways to deal with asynchronous access to mutable data.
We can use isolated and nonisolated keywords to control mutable and immutable states’ access precisely.</p>]]></content><author><name></name></author><category term="ios" /><category term="swift" /><category term="concurrency" /><summary type="html"><![CDATA[By default, each mutable property and method is isolated, which means they can’t be accessed from the external code. If you want to call them, you need to use await keyword. Immutable, constant properties and methods that you mark explicitly with nonisolated keywords can be accessed directly from the external code because they are in a non-isolated state. Here is an example:]]></summary></entry><entry><title type="html">Swift Actors: Prevent data races</title><link href="https://juraskrlec.github.io/ios/swift/concurrency/2022/05/23/swift-actors.html" rel="alternate" type="text/html" title="Swift Actors: Prevent data races" /><published>2022-05-23T00:00:00+00:00</published><updated>2022-05-23T00:00:00+00:00</updated><id>https://juraskrlec.github.io/ios/swift/concurrency/2022/05/23/swift-actors</id><content type="html" xml:base="https://juraskrlec.github.io/ios/swift/concurrency/2022/05/23/swift-actors.html"><![CDATA[<p><a href="https://github.com/apple/swift-evolution/blob/main/proposals/0306-actors.md">Swift Actors</a> are a new type in Swift 5.5. They are part of concurrency changes. The Swift concurrency changes aim to detect data races and prevent all the concurrency bugs that can be tricky to find and fix.
It is essential to understand concurrency, what data races are, and how to tackle them. This article will explore Actors, how they work and how to use them in your project.</p>

<h2 id="what-is-an-actor">What is an Actor?</h2>

<p>Swift has various types that we can use: classes, structs, and enums. Classes are reference types, and structs are value types. 
When you use classes, you must remember that they declare shared mutable states across the program. We can see that that can be tricky when you mix them with concurrency. You must do manual synchronization to avoid data races. 
Data races can occur when two threads simultaneously access and write the same data. Data races often lead to strange and unpredictable behavior in your application.</p>

<p>An Actor is a reference type that protects access to its mutable state from data races. You define an Actor using an <code class="language-plaintext highlighter-rouge">actor</code> keyword:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">GameScore</span> <span class="p">{</span><span class="o">...</span><span class="p">}</span>
</code></pre></div></div>

<p>An Actor can have initializers, methods, properties, and subscripts. It can conform to protocols, extend them and work with generics.</p>

<p>An Actor is a reference type as a class, but there is one thing that you can’t do with an Actor, and that is inheritance:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">BasketballGameScore</span><span class="p">:</span> <span class="kt">GameScore</span> <span class="p">{}</span>
</code></pre></div></div>
<p>You will get this error:</p>
<blockquote>
  <p><em>Actor types do not support inheritance</em></p>
</blockquote>

<h2 id="synchronization---preventing-data-races">Synchronization - preventing Data Races</h2>

<p>We would use various locks to create synchronized access to mutable data to prevent data races. Let’s look at the following example where we want to track the latest game score, for example, basketball. We want to know the latest score, and we want to add the new score:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">class</span> <span class="kt">GameScore</span> <span class="p">{</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">score</span> <span class="o">=</span> <span class="p">[</span><span class="kt">String</span><span class="p">]()</span>
    
    <span class="k">var</span> <span class="nv">latestScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">?</span> <span class="p">{</span>
        <span class="k">return</span> <span class="n">score</span><span class="o">.</span><span class="n">last</span>
    <span class="p">}</span>
    
    <span class="kd">func</span> <span class="nf">add</span><span class="p">(</span><span class="n">_</span> <span class="nv">newScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">score</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="n">newScore</span><span class="p">)</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>This code is fine until we don’t use it in a multi-threaded environment. We can get into a lot of problems if we use this concurrently. To fix this, we can go by the route to do all our reads and writes on the specific DispatchQueue to ensure that all the operations execute in serial order:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">class</span> <span class="kt">GameScore</span> <span class="p">{</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">score</span> <span class="o">=</span> <span class="p">[</span><span class="kt">String</span><span class="p">]()</span>
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">queue</span> <span class="o">=</span> <span class="kt">DispatchQueue</span><span class="p">(</span><span class="nv">label</span><span class="p">:</span> <span class="s">"gameScore.queue"</span><span class="p">)</span>
    
    <span class="k">var</span> <span class="nv">latestScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">?</span> <span class="p">{</span>
        <span class="n">queue</span><span class="o">.</span><span class="n">sync</span> <span class="p">{</span>
            <span class="k">return</span> <span class="n">score</span><span class="o">.</span><span class="n">last</span>
        <span class="p">}</span>
    <span class="p">}</span>
    
    <span class="kd">func</span> <span class="nf">add</span><span class="p">(</span><span class="n">_</span> <span class="nv">newScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">queue</span><span class="o">.</span><span class="n">sync</span> <span class="p">{</span>
            <span class="n">score</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="n">newScore</span><span class="p">)</span>
        <span class="p">}</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>If we want to have a concurrent specific DispatchQueue, we can sync reads and writes by the barrier flag:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">class</span> <span class="kt">GamesScoreQueue</span> <span class="p">{</span>
    <span class="kd">private</span> <span class="k">var</span> <span class="nv">score</span> <span class="o">=</span> <span class="p">[</span><span class="kt">String</span><span class="p">]()</span>
    <span class="kd">private</span> <span class="k">let</span> <span class="nv">queue</span> <span class="o">=</span> <span class="kt">DispatchQueue</span><span class="p">(</span><span class="nv">label</span><span class="p">:</span> <span class="s">"gamesCore.queue.concurrent"</span><span class="p">,</span> <span class="nv">attributes</span><span class="p">:</span> <span class="o">.</span><span class="n">concurrent</span><span class="p">)</span>
    
    <span class="k">var</span> <span class="nv">latestScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">?</span> <span class="p">{</span>
        <span class="n">queue</span><span class="o">.</span><span class="nf">sync</span><span class="p">(</span><span class="nv">flags</span><span class="p">:</span> <span class="o">.</span><span class="n">barrier</span><span class="p">)</span> <span class="p">{</span>
            <span class="k">return</span> <span class="n">score</span><span class="o">.</span><span class="n">last</span>
        <span class="p">}</span>
    <span class="p">}</span>
    
    <span class="kd">func</span> <span class="nf">add</span><span class="p">(</span><span class="n">_</span> <span class="nv">newScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">queue</span><span class="o">.</span><span class="nf">sync</span><span class="p">(</span><span class="nv">flags</span><span class="p">:</span> <span class="o">.</span><span class="n">barrier</span><span class="p">)</span> <span class="p">{</span>
            <span class="n">score</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="n">newScore</span><span class="p">)</span>
        <span class="p">}</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<p>As we can see, we have some manual synchronization to do. To tackle this, we use an Actor. It automatically serializes all synchronized access to its properties and methods. When we refactor our GameScore  class to an Actor, we get this:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="kd">actor</span> <span class="kt">GameScore</span> <span class="p">{</span>
    <span class="k">var</span> <span class="nv">score</span> <span class="o">=</span> <span class="p">[</span><span class="kt">String</span><span class="p">]()</span>
    
    <span class="k">var</span> <span class="nv">latestScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">?</span> <span class="p">{</span>
        <span class="k">return</span> <span class="n">score</span><span class="o">.</span><span class="n">last</span>
    <span class="p">}</span>
    
    <span class="kd">func</span> <span class="nf">add</span><span class="p">(</span><span class="n">_</span> <span class="nv">newScore</span><span class="p">:</span> <span class="kt">String</span><span class="p">)</span> <span class="p">{</span>
        <span class="n">score</span><span class="o">.</span><span class="nf">append</span><span class="p">(</span><span class="n">newScore</span><span class="p">)</span>
    <span class="p">}</span>
<span class="p">}</span>
</code></pre></div></div>

<h3 id="accessing-actors-data">Accessing Actor’s data?</h3>

<p>We don’t know when access is allowed and when another thread will perform access to mutable data, so we can’t just access the mutable property directly. To create asynchronous access, we use await keyword:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="k">let</span> <span class="nv">gameScore</span> <span class="o">=</span> <span class="kt">GameScore</span><span class="p">()</span>
<span class="k">await</span> <span class="n">gameScore</span><span class="o">.</span><span class="n">latestScore</span>
</code></pre></div></div>

<h3 id="data-races-can-still-occur-when-using-actors">Data Races can still occur when using Actors</h3>

<p>Actors can help prevent data races, and they give us an effortless way to synchronize access preventing weird crashes and behaviors. But we need to be careful; they are not a universal solution in a multi-threaded environment where we can forget about data races and race conditions. What if we have two queues, one accessing our score and one storing it:</p>

<div class="language-swift highlighter-rouge"><div class="highlight"><pre class="highlight"><code><span class="n">queueOne</span><span class="o">.</span><span class="k">async</span> <span class="p">{</span>
    <span class="nf">print</span><span class="p">(</span><span class="k">await</span> <span class="n">gamesScore</span><span class="o">.</span><span class="n">latestScore</span><span class="p">)</span>
<span class="p">}</span>

<span class="n">queueTwo</span><span class="o">.</span><span class="k">async</span> <span class="p">{</span>
    <span class="k">await</span> <span class="n">gamesScore</span><span class="o">.</span><span class="nf">add</span><span class="p">(</span><span class="s">"30:32"</span><span class="p">)</span>
<span class="p">}</span>
</code></pre></div></div>

<p>We don’t know which one will fire first, so we can have different behavior and inconsistent game score.</p>

<h2 id="conclusion">Conclusion</h2>

<p>Actors are straightforward and neat ways to deal with asynchronous access to mutable data. They provide exquisite, readable, and one-way solutions to prevent data races. 
In future articles, I’ll cover Actor Isolation, Sendable, and MainActor. So stay tuned!</p>]]></content><author><name></name></author><category term="ios" /><category term="swift" /><category term="concurrency" /><summary type="html"><![CDATA[Swift Actors are a new type in Swift 5.5. They are part of concurrency changes. The Swift concurrency changes aim to detect data races and prevent all the concurrency bugs that can be tricky to find and fix. It is essential to understand concurrency, what data races are, and how to tackle them. This article will explore Actors, how they work and how to use them in your project.]]></summary></entry><entry><title type="html">Welcome to the site!</title><link href="https://juraskrlec.github.io/ios/dev/2022/05/17/welcome-to-site.html" rel="alternate" type="text/html" title="Welcome to the site!" /><published>2022-05-17T00:00:00+00:00</published><updated>2022-05-17T00:00:00+00:00</updated><id>https://juraskrlec.github.io/ios/dev/2022/05/17/welcome-to-site</id><content type="html" xml:base="https://juraskrlec.github.io/ios/dev/2022/05/17/welcome-to-site.html"><![CDATA[<p>Hi, welcome to my site! 
My name is Jura, and I am currently a Staff Engineer at <a href="https://microblink.com">Microblink</a>, specializing in iOS. My work focuses on the SDKs and R&amp;D.</p>

<p>At Microblink, we have developed some fascinating things in the Computer Vision area. Yes, of course, we use Machine Learning :) </p>

<p>On this site, I will focus on the framework side of iOS development, write some tips and tricks, and see where we can go from there. If you have any questions, feel free to contact me! Have a great day!</p>]]></content><author><name></name></author><category term="ios" /><category term="dev" /><summary type="html"><![CDATA[Hi, welcome to my site!  My name is Jura, and I am currently a Staff Engineer at Microblink, specializing in iOS. My work focuses on the SDKs and R&amp;D.]]></summary></entry></feed>