Moreh Unveils Revolutionary AI Inference Technology at AMD Advancing AI 2026
Moreh Unveils Revolutionary AI Inference Technology at AMD Advancing AI 2026
On July 22-23, 2026, in San Francisco, Moreh, an AI infrastructure software company led by CEO Gangwon Jo, took center stage at AMD's flagship annual event, the AMD Advancing AI 2026. During the conference, Moreh demonstrated its cutting-edge distributed inference solution, the MoAI Inference Framework, specifically configured to operate on AMD GPUs.
The highlight of the event was a live demonstration of the GLM-5.1 large language model (LLM), driven by the MoAI Inference Framework, executed on a system featuring 32 AMD Instinct™ MI300X GPUs distributed across four nodes. Attendees had the opportunity to engage with the chatbot functionality firsthand, offering feedback on response speeds and service quality while witnessing performance across a variety of real-world applications.
Unlike typical presentations that merely run an AI model, Moreh's showcase provided essential inference service metrics in real-time. Metrics included GPU utilization rates, Tokens Per Second (TPS), Time To First Token (TTFT), and Time Per Output Token (TPOT). This transparency allowed attendees to assess the inference performance and GPU resource efficiency in a simulated service environment.
Global leaders from the AI industry and enterprise clients, who attended the event, expressed considerable interest in the swift response times and stable performance exhibited by Moreh's technology. They particularly noted the innovative live deployment of the resource-intensive GLM-5.1 model on AMD GPUs at levels matching production-grade services received significant acclaim from visitors.
Moreh's MoAI Inference Framework is touted as the first commercially deployed distributed inference solution tailored for the AMD ecosystem. Its blend of distributed inference and heterogeneous computing technologies aims to significantly lower AI service costs, thereby facilitating wider adoption of AI applications worldwide. This innovation addresses a pressing challenge within the industry—escalating infrastructure and service expenses associated with the growing size of AI models—by delivering a more efficient inference framework.
Gangwon Jo, CEO of Moreh, stated, “This event provided an opportunity for global customers to verify firsthand that top-tier inference performance can be achieved on AMD GPU environments.