Principal Engineer – NPU Compiler & Architecture
NXP Semiconductors · Hyderabad
onsitefull-time10+ years
posted 1d
About the Role We are seeking a highly experienced Compiler / Software Architect to join our NPU Hardware Architecture team and play a key role in the hardware-software co-design of next-generation AI inference accelerators. This is a unique role at the intersection of NPU hardware architecture, AI compilers, performance modeling, and software development. The position will work directly with hardware architects to define, evaluate, and prototype the software stack required to program current and future NPU architectures. Our NPU architecture is closely coupled with the compiler stack. Efficient utilization of the accelerator requires deep understanding of the hardware compute architecture, memory hierarchy, dataflow, scheduling, quantization, instruction set and execution model. The successful candidate will use this understanding to develop compiler and software proof-of-concepts, evaluate architectural proposals against real AI workloads, and influence the design of future NPU hardware. You will work closely with hardware architects, micro-architects, compiler engineers and AI software teams to answer a fundamental question: How should the NPU hardware and software be co-designed to deliver the best performance, power efficiency, programmability and scalability for real-world AI workloads? What You Will Be Responsible For Compiler & Software Architecture for NPU Develop compiler and software proof-of-concepts for current and next-generation NPU architectures. Define software abstractions and compiler flows that efficiently expose NPU hardware capabilities to AI workloads. Develop and evaluate compiler concepts including graph lowering, intermediate representations, operator mapping, scheduling, tiling, fusion, memory planning and code generation. Translate NPU architectural concepts into executable software models and demonstrate their feasibility using representative AI workloads. Develop lightweight compiler/runtime infrastructure to validate new hardware features before production software implementation. Analyze existing compiler limitations and identify architectural changes required to improve programmability and accelerator utilization. Hardware–Software Co-Design Work as an integral member of the hardware architecture team to co-design NPU hardware and software. Analyze how proposed hardware features can be effectively exposed through the compiler and software stack. Provide software-driven feedback on compute architecture, memory hierarchy, data movement, dataflow, scheduling, instruction set and accelerator programmability. Identify hardware features that provide meaningful benefits to real AI workloads and challenge features that add hardware complexity without sufficient software value. Define compiler requirements and software abstractions for new NPU capabilities. Participate in architecture and micro-architecture reviews and influence hardware decisions from a software and workload perspective. Evaluate architectural trade-offs considering performance, power, area, compiler complexity and software scalability. AI Workload & Performance Analysis Analyze representative AI models and workloads to identify compute, memory, bandwidth, scheduling and data-movement bottlenecks. Build software-based performance models and workload prototypes to evaluate architectural concepts. Develop experiments to quantify the impact of proposed hardware features on model performance and accelerator utilization. Investigate issues such as quantization, sparsity, operator fusion, tensor layouts, tiling, data reuse, memory bandwidth and scheduling efficiency. Correlate software/model-level performance with architectural and micro-architectural behavior. Use workload analysis to guide both current-generation optimizations and next-generation NPU architecture. Architecture Prototyping Rapidly prototype software solutions for architectural concepts that may be months or years away from production silicon. Develop functional models, comp