Backed by top-tier venture and strategic investors, we are a stealth mode startup, pioneering system-level innovations for data center-scale inference enabled by an entirely new category of SOC.
Full-time · Onsite · Santa Clara, CA.
CUDA/ROCm/Triton developer to build and optimize high-performance GPU kernels for next-generation AI systems
Full-time · Onsite · Santa Clara, CA.
Drive performance improvements in inference frameworks (vLLM, SGLang, PyTorch) and develop advanced cluster scheduling algorithms to optimize throughput and latency.
Full-time · Onsite · Santa Clara, CA.
Build compiler capabilities from PyTorch through Triton down to CUDA and machine IR on leading-edge accelerators.
Full-time · Onsite · Santa Clara, CA.
Develop performance and/or functional models that guide architecture decisions and accelerate hardware–software co-design.
Full-time · Onsite · Santa Clara, CA.
Build functional and performance models across abstraction levels to guide architecture and accelerate HW-SW co-design.
Full-time · Onsite · Santa Clara, CA.
Define compute blocks and data movement algorithms across all interfaces and compute blocks on a leading-edge SoC.
Full-time · Onsite · Santa Clara, CA.
Drive verification of complex next-generation AI/compute platforms
Full-time · Onsite · Santa Clara, CA.
Implement RTL blocks in close collaboration with architects and micro-architects.
2445 Augustine Dr Ste 150
Santa Clara, CA 95054