Portrait of Reese Kuper

Computer architecture · systems infrastructure · performance

Hi, I'm Reese.

I work at the intersection of computer architecture, systems, and infrastructure - building low-level tools and experiments to measure and optimize complex computing platforms.


Selected Work

Interactive tutorial archive

Intel On-chip Accelerators

Companion site I built for our ISCA 2023 tutorial on Intel QAT, DLB, DSA, and IAA, preserved with the original technical material and resources.

Explore the tutorial site

Quantitative systems research

Data Streaming Accelerator Analysis

A quantitative study of Intel DSA behavior, performance tradeoffs, and practical guidelines for using data-movement acceleration in modern Xeon systems.

GPU scheduling research

Finer-Grained ML Synchronization

YArch work exploring finer-grained synchronization as a way to improve GPU utilization in machine-learning workloads.

Read the workshop paper

Experience

Performance Engineer:

  • Created and supported tools and infrastructure for large-scale sweep experiments in interconnect system analysis.
  • Developed a library around IBM Spectrum LSF for flexible job offloading and better orchestration in HPC clusters.
  • Supported performance model productization via Python scripts to stitch GUI application outputs to model inputs.
  • Dynamically mapped registers between the RTL and performance model for the hardware emulation team.

RTL Design Rotation:

  • Added ASIL-B compliant RTL parity and interface checkers in AXI-Stream blocks.
  • Extended support for new RTL features to existing Python-based Verilog generator tools.

Formal Verification Rotation:

  • Designed a formal testbench for compression and AMU blocks, working with their respective RTL owners.
  • Improved a protocol domain bridge testbench in collaboration with the formal team.
  • Performance analysis on Data Streaming Accelerator (DSA), an on-chip accelerator found on newer Intel Xeon processors.
  • Explored use cases for DSA to take advantage of cache pollution mitigation and higher throughput for memory operations.
  • Submitted patents for improving memory deduplication techniques using DSA
  • Prepared and presented a tutorial website for Intel's on-chip accelerators at ISCA 2023 - earned Intel's internal Department Recognition Award.
  • Investigated the effectiveness of an L1 cache in a CXL device for DLRM offloading.
  • Built a cache simulator to evaluate both hit rates and cache occupancy rates of real-world DLRM memory traces.
  • Analyzed the characteristics of the DLRM data for locality patterns to aid in the design of the memory system.
  • Simulated DRAM and cache designs in Ramulator to obtain bandwidth, hit rate, and other metrics used in analysis.
  • Debugged failing signatures for the CPU memory system testbench.
  • Implemented a new statistical coverage (SCOV) workflow:
    • Used scoreboard listener functions to implement SCOV events for assessing stimulus coverage.
    • Coded a macro to simplify adding SCOV events within the UVM testbench.
    • Developed a Python script to parse regression simulation logs and send results to a MySQL database.
  • Developed Python script to analyze the use of all plusargs within the UVM testbenches.
  • Scripted two Git hooks for file update notifications.
  • Fixed UVM register definition autogeneration for more flexible RAL models.
  • Programmed module for modeling transactions between a master device to interconnect return nodes in SystemC.
  • Formally verified round-robin and LSB-priority arbiters using SystemVerilog assertions.
  • Improved kernel ION memory allocation speeds by ~10%.
  • Analyzed the efficiency of IOVA’s use of caching and compared it with mmap’s gap searching RBTree structure.
  • Created internal Python file tracing tool for parsing Linux RAM dump binaries.
  • Worked towards shifting mmap allocations to use the mempool API.

Education

  • Parallel Computer Architecture
  • Computer Architecture
  • Digital Design and Synthesis
  • Parallel Programming
  • Operating Systems
  • Algorithms
  • Artificial Intelligence
  • OOP and Data Structures
  • Advanced Computer Architecture
  • Computer Microarchitecture
  • System-on-Chip Design
  • Memory and Storage Systems
  • Programming Languages and Compilers

Publications and Patents

A Quantitative Analysis and Guidelines of Data Streaming Accelerator in Modern Intel Xeon Scalable Processors

Reese Kuper, Ipoom Jeong, Yifan Yuan, Ren Wang, Narayan Ranganathan, Nikhil Rao, Jiayu Hu, Sanjay Kumar, Philip Lantz, Nam Sung Kim

[PAPER] International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). April, 2024

[PAPER] Open-access ePrint Archive (arXiv). April, 2023

Efficiently Merging Non-Identical Pages in Kernel Same-Page Merging (KSM) for Efficient and Improved Memory Deduplication and Security

Reese Kuper, Yifan Yuan, Ren Wang

[PATENT] US Patent App 18/369,090. 2024

Method and Apparatus for Batching Pages for a Data Movement Accelerator

Reese Kuper, Yifan Yuan, Ren Wang

[PATENT] US Patent App 18/477,628. 2024

Demystifying CXL Memory with Genuine CXL-Ready Systems and Devices

Yan Sun, Yifan Yuan, Zeduo Yu, Chihun Song, Reese Kuper, Jinghan Huang, Houxiang Ji, Siddharth Agarwal, Jiaqi Lou, Ipoom Jeong, Ren Wang, Jung Ho Ahn, Tianyin Xu, Nam Sung Kim

[PAPER] International Symposium on Microarchitecture (MICRO). October, 2023

[PAPER] Open-access ePrint Archive (arXiv). April, 2023

On-chip Accelerators in 4th Gen Intel® Xeon® Scalable Processors: Features, Performance, Use Cases, and Future!

Reese Kuper, Ipoom Jeong, Yifan Yuan, Jiayu Hu, Ren Wang, Narayan Ranganathan, Nam Sung Kim

[TUTORIAL] International Symposium on Computer Architecture (ISCA). June, 2023

STYX: Exploiting SmartNIC Capability to Reduce Datacenter Memory Tax

Houxiang Ji, Yan Sun, Mark Mansi, Yifan Yuan, Jinghan Huang, Reese Kuper, Michael Swift, Nam Sung Kim

[PAPER] The USENIX Annual Technical Conference (ATC). July, 2023

Improving GPU Utilization in ML Workloads Through Finer-Grained Synchronization

Reese Kuper, Suchita Pati, Matthew D. Sinclair

[PAPER] Young Architect Workshop (YArch). April, 2021


Hobbies

Things I enjoy outside of working with computing systems.

Drawing

I typically draw places I have been to as a way of keeping mementos.

View of Porto

Date drawn
May 2026
Subject
Porto, Portugal from the Ponte Luis 1 Bridge
Materials
Ink on A5 sketchbook

Walking to the Lab at UIUC

Date drawn
May 2025
Subject / source
North Quad towards Beckman Institute at UIUC
Materials
Ink on A5 sketchbook

Second Dorm in Undergrad

Date drawn
May 2022
Subject / source
Second dorm room at UW - Madison
Materials
Ink on A5 sketchbook

First Dorm in Undergrad

Date drawn
January 2018
Subject / source
First dorm room at UW - Madison
Materials
Ink on A5 sketchbook

Baking

Pictures of some of the different recipes that I enjoyed making.


Website made by Reese Kuper