Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

Week one of the Musk v. Altman trial: What it was like in the room

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Two of

Tailoring AI solutions for health care needs

The AI market is full of big promises of grand transformation. Health care is a prime target for those promises, beset as it is by

Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

arXiv:2605.00254v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) architectures have turned LLM serving into a cluster-scale workload in which communication consumes a considerable portion of LLM

AlphaInventory: Evolving White-Box Inventory Policies via Large Language Models with Deployment Guarantees

arXiv:2605.00369v1 Announce Type: cross Abstract: We study how large language models can be used to evolve inventory policies in online, non-stationary environments. Our work is

BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs

arXiv:2605.00422v1 Announce Type: cross Abstract: Large language models (LLMs) have driven major progress in NLP, yet their substantial memory and compute demands still hinder practical

Week one of the Musk v. Altman trial: What it was like in the room

Tailoring AI solutions for health care needs

Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

AlphaInventory: Evolving White-Box Inventory Policies via Large Language Models with Deployment Guarantees

BWLA: Breaking the Barrier of W1AX Post-Training Quantization for LLMs

Rethinking Network Topologies for Cost-Effective Mixture-of-Experts LLM Serving

Subscribe for Updates