Rackline Research
What was just published
Papers are the latest arXiv lists in distributed computing, hardware architecture, and systems, kept only when the title is about compute, power, or the plant. Under each title, a short note says what the paper is for and what the authors conclude, written so a reader outside the lab can follow it. The note uses the abstract and the conclusion. It does not add a result that is not in those two places. Patents are kept only when the title names a data center and the publication is inside the last eighteen months.
Publications
The latest lists, kept only when the title is on the beat
Distributed computing · 2610.08268
DySCo: Dynamic Sharding for Collaborative Edge-Cloud LLM Inference with Depth-Synchronized BatchingJingpo Xu, Paul Joe Maliakel, Ivona Brandic and others
What it is for
Pervasive intelligent applications are increasingly deployed on mobile and Internet of Things (IoT) edge devices.
What follows
Layer-wise edge–cloud large language model inference allows devices with different capabilities to contribute to the same model, but also changes how a shared cloud GPU receives work. DySCo uses resident model shards to support configurable cuts without reloading weights, while keeping each session’s KV caches with the layers that produce them. The single-edge measurements show that moving more layers to the edge does not necessarily reduce latency: the resulting idle interval can increase the lantency of the cloud’s next forward pass.
Distributed computing · 2610.08238
GPU Acceleration of Awkward Arrays: Using Python cuda.computeMaksym Naumchyk, Ianna Osborne
What it is for
Awkward Array is a widely used library in high-energy physics (HEP) for representing and manipulating nested, variable-length data in Python. The authors present performance studies that demonstrate improvements over previously reported eager GPU execution strategies for representative HEP analysis patterns.
What follows
The conclusion was not in the HTML copy, so this note stops at what the paper is for.
Distributed computing · 2610.07504
Mosaic: GPU Sharing with Latency Guarantees through Kernel-Level Interference PredictionFoteini Strati, Ethan Graham, Leo Stephan and others
What it is for
GPUs are increasingly in demand for AI workloads, yet often remain substantially underutilized, motivating workload colocation. The authors present Mosaic, a kernel-level interference predictor that explicitly models these mechanisms using a combination of analytical and lightweight learned models.
What follows
The authors presented Mosaic, an accurate and lightweight kernel-level predictor for GPU interference. Mosaic combines analytical models for thread block scheduling, warp scheduling, and compute pipeline contention with xgboost models for the memory subsystem, achieving up to an order of magnitude lower p50 and p95 error than prior predictors while keeping per-prediction latency in the microsecond range. Using Mosaic, they built MosaicSched, a GPU scheduler for workload colocation.
Distributed computing · 2610.08378
Lachesis: Lifetime-Aware KV Cache Placement for Agent Serving across HBM and High-Bandwidth FlashJaehoon Yang, Jeongmin Lee, Haneul Park and others
What it is for
Large language model serving is increasingly dominated by agentic workloads, in which agents and their sub-agents accumulate context as KV cache across many requests, consuming substantial memory.
What follows
The authors find that the context an agent session accumulates divides at harness-defined context boundaries into context segments whose lifetimes differ along three axes, temporal, structural and inter-worker, and that each lifetime is knowable before its segment is written. Lachesis places each segment on the tier its lifetime implies, along a deterministic path that reads the segment class the harness fixes and a predictive path that ranks worker segments at spawn.
Distributed computing · 2610.07782
Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can TellHochan Son, Kyungdoe Han, Jaehan Koh and others
What it is for
Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. The authors measure both on one three-tier agent architecture.
What follows
The conclusion was not in the HTML copy, so this note stops at what the paper is for.
Distributed computing · 2610.07219
Cascadia: Resident 975B MoE Inference on Eleven AI PCsTate Berenbaum (Not Community Labs Inc.), Matias Parij (Not Community Labs Inc.), Muthaiah Venkatachalam (Intel Corporation)
What it is for
Mixture-of-experts models make nearly trillion-parameter capacity accessible with sparse per-token computation, provided that the serving system can distribute the weights and coordinate their execution.
What follows
Cascadia demonstrates resident inference of Inkling’s 975B MoE decoder on eleven Panther Lake AI PCs. Its custom engine preserves Inkling’s routing rules, constructs compressed graphs for OpenVINO’s fused iGPU primitives, manages numerical range across FP16 and FP32, and unifies dense and sparse blocks within resident layer shards. Stream-oriented coordination connects those shards across machines.
Distributed computing · 2610.07126
CUDA-MPQS: A GPU-Resident Self-Initializing Quadratic Sieve, and the Factorization of RSA-155Fabian Januszewski, Christoph Heinrichs
What it is for
The quadratic sieve (QS) is an irregular, cache-hostile, branch-heavy integer factorization algorithm; prior GPU work accelerates individual stages of it.
What follows
Three extensions are quantified by their own results. The authors do not rank it first: the number field sieve overtakes the quadratic sieve near 100–110 digits regardless of variant and the gap widens super-polynomially (§V-C offers one same-machine data point; on GPUs, concurrent work reports a 155-digit GNFS in about 1.2 GPU-days ), so multiple large primes would pay exactly where the quadratic sieve is the wrong algorithm, and nothing here bears on deployed parameters. GPU sievers are unlikely to be near their ceiling: v1.0.8’s binary alone cuts RSA-100 wall-clock by 19–21% over v1.0.7 on two devices, and 18% and 36% on two more across runs (Table I); the lattice siever’s richer structure may hide more.
Distributed computing · 2610.07094
Evaluating Inference Compute for Generative AI: A Framework for Enterprise WorkloadsAbbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
What it is for
Large language model deployment is shifting from single-turn completion to agentic trajectories in which a model plans, calls tools, reads results and reasons at test time before acting.
What follows
Agentic AI turns inference from batchable completions into sequential trajectories, moving the hardware bottleneck from throughput to single-stream decode latency-a quantity governed by weight-tier bandwidth, not peak compute. Reliability compounds exponentially in trajectory length, so determinism and tail latency are performance variables. The evaluation question is therefore no longer “GPU or specialised?” but which prefill/decode pair, through which channel, at what utilisation-measured as goodput at an agentic SLO and cost per successful episode, with paired statistics, attested vendor evidence, and break-even conditions on capital ratio, utilisation and adoption lead time.
Distributed computing · 2610.06718
One Global Beam Across Many GPUs: High-Throughput Beam Search at Billion-Record Frontier ScaleIvan Litvak
What it is for
Beam search repeatedly makes many children, removes duplicates, and keeps the best . The authors show how many GPUs can perform these steps as one search even when the retained set does not fit on one device.
What follows
The conditional abstract beam-selection result can be summarized in one line: for the same complete candidate set, quantized scores, Hash128 key equivalence, payload order, fully specified rank/physical-buffer/local-slot tie order, distinct-key threshold witnesses, global one-per-key reduction, race-free routing, and complete processing without semantic local caps. The equality is selected membership under the same realized physical layout, not a canonical serialization order or schedule-independent result from GPU and shard counts alone. Full-state equality additionally requires no relevant Hash128 collision.
Distributed computing · 2610.06622
GPU-Initiated Discrete Simulated Bifurcation: Low-Latency Requests and Streaming Dense CouplingsYaocheng Chen
What it is for
GPU-based optimization faces two communication bottlenecks: coordinating frequent requests and delivering dense models that exceed device memory. The authors present a discrete simulated bifurcation (dSB) architecture that addresses both through NVIDIA DOCA GPUNetIO.
What follows
The authors implemented dSB as a GPU-initiated request service and a solver for streamed dense couplings. The request service keeps receive, solve, and reply operations on the GPU, using one worker block per request and batched transmission. It requires no dedicated server CPU data-path core.
Distributed computing · 2610.05748
MOLT: A Fine-Grained GPU Memory Sharing System for LLM Serving with Opportunistic Fine-TuningJaehoon Yang, Yongbeom Kim, Hojoon Kim and others
What it is for
Large language model serving scales its replica count with the request load, yet GPU memory still stands idle inside the replicas.
What follows
Autoscaling leaves GPU memory idle inside large language model serving replicas, because it cannot follow a KV cache that changes within seconds and removes a replica only as a whole. Existing colocation systems either keep the tuning memory resident or let inference reclaim it only in whole training samples, discarding the running tuning step each time. The authors present MOLT, a fine-grained memory sharing system for colocating large language model inference and tuning.
Distributed computing · 2610.05305
Characterizing Parallelism Strategies in LLM Inference: Fundamental Compute-Communication Trade-offsJavad Mirzaei, Jeebak Mitra
What it is for
Large Language Model inference has become the dominant workload in modern AI systems, requiring serving infrastructures to maximize throughput while meeting strict latency Service-Level Objectives (SLOs).
What follows
In this paper, they presented a unified analytical framework for modeling distributed large language model inference by decomposing end-to-end latency into computation, inter-GPU communication, and pipeline bubble overhead. The authors derived analytical models for TP collective communication, PP point-to-point communication, and pipeline utilization, capturing their dependence on batch size, sequence length, hidden dimension, message size, and parallelism degree. Unlike existing empirical studies, their framework provides a principled explanation of how these factors jointly determine inference performance across both the prefill and decoding phases.
Patents
Titles that name a data center
The patent pull did not return a recent document whose title names a data center. Body-text matches are discarded. An empty list here is the honest result, not a failure to look.