Claiming the $100 cloud credit from the AMD AI Developer Program involves a two-stage signup, a confusing DigitalOcean email loop, and non-obvious Credit balances. Here is the exact path through the UI.
Tested on EVO-X2 (gfx1151) under Windows + ROCm: the official b10666 binary runs it at pp512 159.25 / tg128 23.26, and the earlier crash was my own workaround flag.
Tested on GMKtec EVO-X2 (Ryzen AI Max+ 395): Q8_0 + MTP beats a 4-bit M1 Max at 22 tok/s, thinking burns 32,712 chars before any HTML, and the NSFW refusal line moves.
Tested local Wan video gen on a Radeon 8060S (Strix Halo, 48GB UMA, Windows). ZLUDA can't run stock PyTorch; AMD's TheRock gfx1151 wheel gives native ROCm. FastWan 1.3B in 4min, Wan 14B I2V in 13.6min — VAE decode and 16GB-RAM Segfaults are the real limits.
Based on EE Times' interview with AMD AI Software VP Anush Elangovan, we assess the ROCm vs CUDA ecosystem gap. Includes hands-on experience with ROCm breaking four times on Strix Halo, plus practical guidance on choosing between NVIDIA, AMD, and Apple Silicon.
Benchmarking NII's LLM-jp-4-32B-A3B-thinking on EVO-X2 (Ryzen AI Max+ 395) with ROCm. 62.9 t/s vs Qwen3.5-35B-A3B's 44.7 t/s. Covers thinking control issues, KV cache trade-offs, knowledge cutoff, Japanese quality comparisons, code generation tests, and training data composition.
Lemonade is AMD's open-source local AI server that manages multiple backends like llama.cpp and FastFlowLM across GPU/NPU/CPU, serving text, image, and audio generation through an OpenAI-compatible API.
Only 10 of 40 layers use KV cache, so raising llama-server ctx-size from 4096 to 65536 cost 800MB VRAM and no throughput. Measured on Ryzen AI Max+ 395.
After updating to AMD Software 26.3.1 on a GMKtec EVO-X2 (Ryzen AI Max+ 395), Vulkan backend fails to allocate device memory properly and falls back to CPU. Investigation and workaround by changing BIOS VRAM allocation from 48GB/16GB to 32GB/32GB.
Hands-on test of huihui-ai Qwen 3.5 abliterated models in Ollama: garbage-token failures, GLM-4.7-Flash chat-template breakage, and why the official model with thinking disabled worked better.