Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

mattpd

@mattpd@mastodon.social
mastodon 4.8.0-nightly.2026-10-06
  • Open on mastodon.social

https://github.com/MattPD
https://twitter.com/matt_dz

0 Followers
0 Following
11 Posts
Joined April 20, 2017
Open post
mattpd @mattpd@mastodon.social
· 1mo ago

mold: A Massively Parallel Linker
https://arxiv.org/abs/2608.23228
Rui Ueyama
International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) 2027

mold: A Massively Parallel Linker
arXiv.org

mold: A Massively Parallel Linker

Linking is a critical step in the software build process that combines compiled object files into a single executable or shared library. Despite decades of engineering effort, link times remain a significant bottleneck in the edit-compile-debug cycle, particularly for large C++ programs. Existing linkers exploit limited parallelism, leaving most CPU cores idle during linking. We present mold, a Unix/Linux linker that applies data parallelism systematically across the entire linking pipeline. We

15
0
8
0
Open post
mattpd @mattpd@mastodon.social
· 1mo ago

Lost Bytes At The Crossroads Between User- And Kernel-Level Memory Allocation
https://www.ibr.cs.tu-bs.de/vss/Publications/2026/fistanto_26_lost_bytes.pdf
Pasha Fistanto, Sören Tempel, and Christian Dietrich
14th Workshop on Programming Languages and Operating Systems (PLOS) 2026

ibr.cs.tu-bs.de
1
0
1
0
Open post
mattpd @mattpd@mastodon.social
· 7mo ago

When AI Writes the World’s Software, Who Verifies It?

https://leodemoura.github.io/blog/2026/02/28/when-ai-writes-the-worlds-software.html

"Writing a specification forces clear thinking about what a system must do, what invariants it must maintain, what can go wrong. This is where the real engineering work has always lived. Implementation just used to be louder."

leodemoura.github.io
8
0
4
0
Open post
mattpd @mattpd@mastodon.social
· 5mo ago
Replying to

@TomF@mastodon.gamedev.place On that note, also like John Ousterhout's "Always measure one level deeper", https://al.radbox.org/doi/10.1145/3213770

"If you want to understand the performance of a system at a particular level, you must measure not just that level but also the next level deeper. That is, measure the underlying factors that contribute to the performance at the higher level."

al.radbox.org

Always measure one level deeper

Performance measurements often go wrong, reporting surface-level results that are more marketing than science.

2
0
0
0
Open post
mattpd @mattpd@mastodon.social
· 9mo ago
Replying to
@neilhenning@mastodon.gamedev.place @nikic@mastodon.social Somewhat related (a different ecosystem, albeit MLIR-based) but this looks like an interesting idea: Triton Extensions: a framework for developing and building Triton compiler extensions https://github.com/triton-lang/triton-ext/ Presented a few days ago: https://www.youtube.com/watch?v=JnFFwBB6Dhk&t=781s Allows to develop custom passes, dialects, backends, and language extensions downstream without having to fork the upstream compiler codebase.

Triton Community Meetup 20260107

4
1
0
0
Open post
mattpd @mattpd@mastodon.social
· 6mo ago
Replying to
@glennklockwood BTW, these findings look interesting: https://specdecode-bench.github.io/
specdecode-bench.github.io

Speculative Decoding: Performance or Illusion? | SpecDecode-Bench

The first systematic study of speculative decoding on a production-grade inference engine (vLLM), covering multiple SD variants across diverse workloads, model scales, and batch sizes.

2
0
0
1
Open post
mattpd @mattpd@mastodon.social
· 6mo ago
Replying to
@adrian@discuss.systems Thanks! BTW, am I the only one getting "Abstract: Data missing." with https://al.radbox.org/doi/10.1145/3779212.3790241 It does show up on https://dl.acm.org/doi/10.1145/3779212.3790241 FWIW
al.radbox.org

Triton-Sanitizer: A Fast and Device-Agnostic Memory Sanitizer for Triton with Rich Diagnostic Context

Memory access errors remain one of the most pervasive bugs in GPU programming. Existing GPU sanitizers such as compute-sanitizer detect memory access errors by instrumenting every memory instruction in low-level IRs or binaries, which imposes high overhead and provides minimal memory access error diagnostic context for fixing problems. We present Triton-Sanitizer, the first device-agnostic memory sanitizer designed for Triton, a domain-specific language for developing portable, efficient GPU ker

1
4
0
0
Open post
mattpd @mattpd@mastodon.social
· 6mo ago
Replying to
@adrian@discuss.systems Ah, not just me, then! I'm blaming them wild ASPLOS folks taking ideas like the Asemantic Code Revolution too far! Thanks for the service in any case, this is a vast improvement over the ACM DL "experience" :-)
1
2
0
0
Open post
mattpd @mattpd@mastodon.social
· 6mo ago
Replying to
@glennklockwood See also Speculative Speculative Decoding, https://arxiv.org/abs/2603.03251 https://github.com/tanishqkumar/ssd
Speculative Speculative Decoding
arXiv.org

Speculative Speculative Decoding

Autoregressive decoding is bottlenecked by its sequential nature. Speculative decoding has become a standard way to accelerate inference by using a fast draft model to predict upcoming tokens from a slower target model, and then verifying them in parallel with a single target model forward pass. However, speculative decoding itself relies on a sequential dependence between speculation and verification. We introduce speculative speculative decoding (SSD) to parallelize these operations. While a v

1
2
0
0
Open post
mattpd @mattpd@mastodon.social
· 1mo ago

Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
https://arxiv.org/abs/2608.22602
Accel-Sim 2.0: Validated GPU Simulation with full Hopper support
https://github.com/accel-sim/accel-sim-framework

Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era
arXiv.org

Architecting the Next Generation of Asynchronous, Distributed GPUs for the AI Era

The rapid evolution of machine learning workloads has fundamentally transformed GPU hardware, driving architectures toward Multi-Chip Module (MCM) topologies, asynchronous execution primitives, and persistent, multi-phase kernel behaviors. Despite these shifts, cycle-level simulation infrastructure has lagged behind, lacking the native capability to model the physical non-uniformity of modern GPUs alongside the massive scale of state-of-the-art AI workloads. To bridge this gap, we present a cycl

0
0
0
0
Open post
mattpd @mattpd@mastodon.social
· 6mo ago
Replying to
@adrian@discuss.systems That works, thanks!
0
0
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 00:12:27 UTC