Senior engineer on Amazon's Neuron SDK team (Annapurna Labs) building and tuning distributed LLM inference for AWS Trainium/Inferentia ML accelerators — designing high-performance kernels, profiling bottlenecks, and optimizing across the stack from PyTorch/JAX down to hardware. Core tech: Python, C++, PyTorch, CUDA/Triton-style kernel work.
Sign in to see your match