<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://jesuisalexjamet.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://jesuisalexjamet.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-08-17T00:26:02+00:00</updated><id>https://jesuisalexjamet.github.io/feed.xml</id><title type="html">Alexandre V. Jamet, PhD</title><subtitle>Alexandre V. Jamet, PhD, is an AI4S Fellow at BSC researching memory systems, branch prediction, and simulation. Explore his work on IP-CaT, TLP, and CPU microarchitecture. </subtitle><entry><title type="html">Co-Designing Cache and TLB Management for High-Performance Instruction Prefetching</title><link href="https://jesuisalexjamet.github.io/blog/2026/ip-cat/" rel="alternate" type="text/html" title="Co-Designing Cache and TLB Management for High-Performance Instruction Prefetching"/><published>2026-04-26T00:00:00+00:00</published><updated>2026-04-26T00:00:00+00:00</updated><id>https://jesuisalexjamet.github.io/blog/2026/ip-cat</id><content type="html" xml:base="https://jesuisalexjamet.github.io/blog/2026/ip-cat/"><![CDATA[<p>The continuous scaling of scale-out cloud applications, microservice architectures, and modern database management systems has profoundly impacted processor microarchitecture. Unlike traditional loop-heavy computational kernels, contemporary enterprise server workloads are characterized by sprawling, multi-megabyte instruction footprints. This massive expansion of the instruction working set routinely exceeds the capacity of standard Level-1 Instruction (L1I) caches, resulting in frequent front-end starvation and establishing instruction supply as a primary execution bottleneck in modern Out-of-Order (OoO) architectures.</p> <p>To mitigate this bottleneck, substantial research has been dedicated to developing sophisticated state-of-the-art L1I prefetchers operating within the virtual address space. However, as demonstrated in our recent <strong>ISCA 2026</strong> publication, the efficacy of even the most advanced prefetching algorithms is fundamentally constrained by underlying hardware management structures.</p> <p>In our paper, <strong>“Enhancing Instruction Prefetching via Cache and TLB Management”</strong>, we characterize a critical microarchitectural “double bottleneck”: the decoupling of prefetch address generation from the virtual-to-physical translation path and the unmanaged insertion of instruction streams into the lower cache hierarchy. To overcome these limitations, we propose <strong>IP-CaT</strong> (Instruction Prefetch Centric Cache and TLB Management), an integrated, low-overhead microarchitectural framework that co-designs Translation Lookaside Buffer (TLB) and cache management policies around the unique behaviors of instruction streams <a class="citation" href="#jamet_ip_cat_isca_2026">[1]</a>.</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/blog/ipcat_cover_description-480.webp 480w,/assets/img/blog/ipcat_cover_description-800.webp 800w,/assets/img/blog/ipcat_cover_description-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/blog/ipcat_cover_description.png" class="img-fluid rounded z-depth-1 mx-auto d-block" width="100%" height="auto" title="High-level description of IP-CaT." loading="lazy" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> <figcaption class="caption">High-level description of IP-CaT.</figcaption> </figure> <hr/> <h2 id="unveiling-the-microarchitectural-bottlenecks">Unveiling the Microarchitectural Bottlenecks</h2> <p>While contemporary prefetchers (such as EPI, FNL+MMA, and Barça) exhibit high accuracy in predicting future instruction stream trajectories, their real-world execution gains are heavily suppressed by two main microarchitectural phenomena:</p> <h3 id="1-the-translation-latency-barrier-the-virtual-memory-wall">1. The Translation Latency Barrier (The Virtual Memory Wall)</h3> <p>Because sophisticated L1I prefetchers operate utilizing virtual addresses, every generated prefetch line request must undergo address translation before interacting with the physical cache hierarchy. Due to the expansive nature of server code, prefetch sequences frequently cross 4KB virtual page boundaries. When a prefetch request references a page not currently mapped in the iTLB, nor in the sTLB, it triggers a translation miss, forcing an invocation of the page-table walker.</p> <p>Consequently, the translation latency often exceeds the prefetch lead time. By the time the physical address is resolved, the demand fetch from the execution pipeline has already arrived at the corresponding Instruction Pointer (IP), rendering the prefetch <em>late</em> and structurally useless. Furthermore, injecting these speculative page walks directly into the primary Translation Lookaside Buffer hierarchy risks evicting critical, highly reusable demand mappings, thereby degrading overall translation throughput.</p> <h3 id="2-the-cache-pollution-crisis">2. The Cache Pollution Crisis</h3> <p>Instruction streams fundamentally differ from data streams regarding their temporal reuse characteristics. To effectively mask deep memory latencies, virtual prefetchers aggressively inject a high volume of blocks into the Level-2 Cache (L2C). However, instructions typically exhibit highly volatile reuse behavior; a substantial portion of prefetched lines are consumed exactly once by the hardware front-end and never referenced again.</p> <p>Standard L2C insertion and replacement policies (e.g., standard LRU or basic pseudo-LRU variations) fail to differentiate between regular data streams and highly transient instruction prefetch streams. As a result, these “Dead-on-Arrival” (DoA) instruction cache lines reside in the L2 allocation sets for extended periods, causing severe cache pollution by prematurely evicting highly reusable data and instruction blocks, which ultimately suppresses the system’s overall Instructions Per Cycle (IPC).</p> <blockquote> <p><strong>Core Thesis:</strong> Maximizing next-generation front-end performance requires shifting focus from prediction accuracy to structural enablement—ensuring that the underlying cache and memory translation hierarchies are explicitly optimized to handle aggressive prefetch streams.</p> </blockquote> <hr/> <h2 id="the-ip-cat-architectural-framework">The IP-CaT Architectural Framework</h2> <p>To address these coupled challenges, <strong>IP-CaT</strong> introduces two low-overhead, specialized hardware mechanisms designed to eliminate translation delays and cache degradation simultaneously.</p> <div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>+-------------------------------------------------------------------+
|                        IP-CaT Framework Architecture              |
+-------------------------------------------------------------------+
|                                                                   |
|             [ Virtual Address Prefetch Stream ]                   |
|                              |                                    |
|                              v                                    |
|     +------------------------+--------+-----------------------+   |
|     |  tPB (Translation      |        |  TIPRP (Trimodal      |   |
|     |  Prefetch Buffer)      |        |  L2C Replacement)     |   |
|     +------------------------+        +-----------------------+   |
|     | Intercepts page-cross  |        | Decision-tree driven; |   |
|     | prefetch translations; |        | Identifies &amp; filters  |   |
|     | decouples sTLB path.   |        | DoA instruction blocks|   |
|     +------------------------+        +-----------------------+   |
|                              |                                |   |
+------------------------------|--------------------------------|---+
v                                v
[ Timely Prefetch Execution ]     [ Optimized L2C Capacity ]
</code></pre></div></div> <p><em>(Figure 1: Architectural block diagram illustrating the dual-component execution flow of the IP-CaT framework.)</em></p> <h3 id="1-decoupling-translations-via-the-tpb">1. Decoupling Translations via the tPB</h3> <p>To insulate the critical demand translation path from speculative prefetch overheads, we introduce the <strong>tPB (translation Prefetch Buffer)</strong>. The tPB functions as a small, highly specialized cache structure situated parallel to the secondary TLB (sTLB) layer.</p> <p>When an L1I prefetch stream detects an impending page-boundary crossing, the tPB eagerly intercepts the virtual address, coordinates with the hardware page-table walker, and caches the resulting translation mapping locally. By isolating these prefetch-induced translations within the tPB, IP-CaT prevents the pollution of the main TLB hierarchy. When the pipeline subsequently generates the corresponding demand fetch, the translation is retrieved instantly from the tPB with near-zero latency, effectively transforming late prefetches into highly effective, on-time hits.</p> <h3 id="2-adaptive-cache-insulation-via-tiprp">2. Adaptive Cache Insulation via TIPRP</h3> <p>To mitigate L2C capacity degradation, we propose the <strong>TIPRP (Trimodal Instruction Prefetch Replacement Policy)</strong>. Traditional cache partitioning or set-dueling frameworks adapt too slowly to the rapid, phase-driven changes characteristic of modern instruction streams. TIPRP replaces these mechanisms with a highly efficient, hardware-implemented <strong>decision tree classifier</strong>.</p> <p>TIPRP monitors the real-time reuse characteristics of incoming blocks, dynamically evaluating the likelihood of an instruction prefetch being Dead-on-Arrival. Based on this real-time telemetry, the policy dynamically shifts each cache set between three specialized sub-policies:</p> <ul> <li><strong>Aggressive Demotion:</strong> Inserting speculative lines with minimal dead-time residency.</li> <li><strong>Instant Eviction:</strong> Bypassing or immediately marking low-confidence prefetch blocks for replacement.</li> <li><strong>Preservation Mode:</strong> Protecting instruction streams that exhibit dense temporal or spatial locality.</li> </ul> <p>By executing this trimodal classification at hardware speeds, TIPRP effectively purges dead prefetch blocks before they can displace critical working sets.</p> <hr/> <h2 id="quantitative-evaluation-and-performance-impact">Quantitative Evaluation and Performance Impact</h2> <p>We evaluated the architectural efficacy of the IP-CaT framework using cycle-accurate simulation across an extensive corporate benchmarking suite encompassing <strong>105 complex server workloads</strong>. The framework was co-evaluated alongside several state-of-the-art instruction prefetchers and compared against contemporary cache and TLB management mechanisms, including Morrigan, CHiRP, Emissary, Mockingjay, and SHiP++.</p> <p>Our experimental results demonstrate that integrating IP-CaT with an advanced Edge Prefetch Indexing (EPI) configuration yields a <strong>6.1% geometric mean speedup</strong> in overall system performance.</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/blog/ISCA_2026_REBUTTALS_all_pref_speedups.svg" sizes="95vw"/> <img src="/assets/img/blog/ISCA_2026_REBUTTALS_all_pref_speedups.svg" class="img-fluid rounded z-depth-1" width="100%" height="auto" title="IP-CaT Empirical Performance Evaluation" loading="lazy" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> </figure> <p><em>(Figure 2: Empirical performance analysis displaying normalized speedups across diverse enterprise server workloads compared to state-of-the-art baseline models including Mockingjay and Morrigan).</em></p> <p>Detailed microarchitectural analysis highlights two primary vectors of improvement:</p> <ul> <li><strong>Translation Latency Mitigation:</strong> The implementation of the tPB successfully counteracts over 80% of the execution penalties historically introduced by page-boundary prefetch serialization.</li> <li><strong>Cache Hit Rate Enhancement:</strong> By filtering out transient instruction streams, TIPRP significantly optimizes the L2C allocation space, minimizing dead block residency without necessitating large tracking structures or high power overheads.</li> </ul> <hr/> <h2 id="conclusion">Conclusion</h2> <p>For decades, microprocessor design paradigms have treated prefetching, virtual memory translation, and cache replacement policies as orthogonal research domains. However, as the instruction footprints of modern cloud-scale software scale past physical hardware boundaries, these isolated silos introduce profound performance liabilities. IP-CaT demonstrates that an integrated, co-designed management layer can unlock latent execution potential within existing execution pipelines.</p> <p>For a comprehensive review of our microarchitectural methodology, hardware overhead analysis, and workload breakdowns, we invite you to read our full preprint on arXiv or engage with our research team directly at ISCA 2026.</p>]]></content><author><name></name></author><category term="computer-architecture"/><category term="microarchitecture"/><category term="cpu"/><category term="front-end"/><category term="instruction-prefetching"/><category term="tlb-management"/><category term="virtual-memory"/><summary type="html"><![CDATA[An architectural deep-dive into the IP-CaT framework published at ISCA 2026, addressing the translation and pollution bottlenecks of instruction streams.]]></summary></entry><entry><title type="html">The Early Shape of a Long Project</title><link href="https://jesuisalexjamet.github.io/blog/2026/the-early-shape-of-a-long-project/" rel="alternate" type="text/html" title="The Early Shape of a Long Project"/><published>2026-04-18T00:00:00+00:00</published><updated>2026-04-18T00:00:00+00:00</updated><id>https://jesuisalexjamet.github.io/blog/2026/the-early-shape-of-a-long-project</id><content type="html" xml:base="https://jesuisalexjamet.github.io/blog/2026/the-early-shape-of-a-long-project/"><![CDATA[<figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/alessandro-erbetta-8oYPewvmhnY-unsplash-480.webp 480w,/assets/img/alessandro-erbetta-8oYPewvmhnY-unsplash-800.webp 800w,/assets/img/alessandro-erbetta-8oYPewvmhnY-unsplash-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/alessandro-erbetta-8oYPewvmhnY-unsplash.jpg" class="img-fluid rounded z-depth-1 mx-auto d-block" width="65%" height="auto" title="The Horizon of Microarchitecture" loading="lazy" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> <figcaption class="caption">The path ahead isn’t always clear — but it’s worth walking.</figcaption> </figure> <h2 id="-a-new-chapter">🎓 A New Chapter</h2> <p>It’s been a few months since I finished my PhD — in September 2024, <em>cum laude</em>, and with a lot of relief and a little pride. I still remember the moment: walking out of the defense room, heart pounding, and seeing my mom waiting with the biggest smile — and a hug that said everything.</p> <figure> <picture> <source class="responsive-img-srcset" srcset="/assets/img/me-and-mom-phd-day-480.webp 480w,/assets/img/me-and-mom-phd-day-800.webp 800w,/assets/img/me-and-mom-phd-day-1400.webp 1400w," type="image/webp" sizes="95vw"/> <img src="/assets/img/me-and-mom-phd-day.jpeg" class="img-fluid rounded z-depth-1 mx-auto d-block" width="65%" height="auto" title="The day I became Dr. Alexandre V. Jamet." loading="lazy" onerror="this.onerror=null; $('.responsive-img-srcset').remove();"/> </picture> <figcaption class="caption">My mom, my first supporter — and my proudest cheerleader.</figcaption> </figure> <p>That moment wasn’t just about me. It was about all the late nights, the coffee-fueled coding sessions, the doubts, and the quiet encouragement that kept me going. And now? I’m ready to take that same energy — and curiosity — into something new.</p> <p>I’d spent years deep in cache hierarchies, memory systems, and the usual suspects of high-performance computing. But I’ve always been drawn to the hard, the unexplored, the things that make you ask: <em>What if we did this differently?</em></p> <p>So I asked myself: What if I turned my attention to the front-end of a processor — the part that fetches and decodes instructions? It’s not the flashiest area. It’s often seen as a place of small, incremental improvements. But what if that’s exactly where a big shift could happen?</p> <h2 id="-why-the-front-end">🤔 Why the Front-End?</h2> <p>For decades, the front-end — the fetch, decode, and branch prediction stages — has been based on the humble instruction as its base work unit. It’s where the CPU reads the code, but not where the <em>meaning</em> of the code is fully understood.</p> <p>We’ve optimized branch prediction with <strong>2-bit counters</strong>, <strong>global history tables</strong>, and even <strong>machine learning</strong> — yet we still treat the instruction as a black box. A signal. A byte sequence.</p> <p>But what if we stopped treating it that way?</p> <p>What if we looked at the <strong>semantic structure</strong> of the program — the loops, the function calls, the control flow patterns — and used that to guide the front-end?</p> <p>That’s the idea I’m starting to explore.</p> <h2 id="-a-glimpse-into-the-work-ahead">🔍 A Glimpse Into the Work Ahead</h2> <p>Over the next few months, I’ll be sharing some early thoughts and experiments — and yes, I know this is a long-term project. But I think it’s worth it.</p> <p>Here’s what I’m thinking about:</p> <ol> <li><strong>Semantic-Aware Branch Prediction</strong> — Instead of just tracking history, what if we used program-level patterns — like loop boundaries or function call sites — to predict whether a branch will be taken?</li> <li><strong>Semantic-Aware Branch Target Prediction</strong> — Can we use the <em>context</em> of a branch — such as the function it’s in — to better predict where it will jump?</li> <li><strong>Rethinking the Decode Stage</strong> — What if we didn’t decode every instruction the same way? Could we skip or simplify decoding for predictable patterns?</li> </ol> <p>These aren’t just ideas — they’re grounded in real data. I’ve already seen <strong>up to a 12% reduction in Branch MPKI</strong> on cloud workloads using non-ML approaches that leverage semantic insights. That’s not a small number — especially when you consider how much performance and energy is lost to mispredictions.</p> <h2 id="-why-this-matters">🧠 Why This Matters</h2> <p>Modern workloads — especially in the cloud — are becoming more complex. We have microservices, containers, serverless functions, and dynamic execution patterns. And yet, our CPUs are still built on assumptions from the 1990s.</p> <p>We’re asking: <em>What if we could design front-ends that are not just faster, but smarter?</em></p> <p>Smarter in the sense that they understand the <em>intent</em> behind the code — not just the bits.</p> <p>This isn’t about replacing ML — it’s about <strong>complementing it</strong>. Or even <strong>replacing it</strong> in some cases, where simplicity, predictability, and energy efficiency matter more than marginal gains.</p> <h2 id="-a-long-term-vision">🌱 A Long-Term Vision</h2> <p>I don’t expect to have all the answers by next month. But I do have a plan:</p> <ul> <li>Start with <strong>small, focused experiments</strong> — testing semantic features on real workloads.</li> <li>Build <strong>prototypes</strong> — not just simulations, but actual hardware-aware models.</li> <li>Share <strong>early results</strong> — even if they’re imperfect — because progress is more valuable than perfection.</li> </ul> <p>And I’d love to have you along for the ride.</p> <p>If you’ve ever wondered whether we’re designing CPUs the right way, or if there’s a better way to think about performance and energy, I’d love to have you along for the ride.</p> <p>Stay tuned. 🚀</p>]]></content><author><name></name></author><category term="computer-architecture"/><category term="microarchitecture"/><category term="cpu"/><category term="front-end"/><category term="branch-prediction"/><category term="hardware-design"/><category term="semantic-analysis"/><summary type="html"><![CDATA[A personal journey into rethinking CPU front-ends — and why I think it's time to look beyond instructions.]]></summary></entry></feed>