Multimodal AI - models that process and reason across text, images, audio, and documents simultaneously - has moved from a differentiating feature to a baseline expectation for frontier models by mid-2026. Both GPT-5 and Gemini Ultra handle multimodal inputs. The question is which one handles them better, for which specific...