Tag: Multimodal
Posts with this tag
- Holo4: One Agent Across GUIs, Code, and Tools—with Different Licenses
A closer look at H Company’s cross-interface agent and long-horizon harness, its benchmark claims, public traces, and the licensing split between checkpoints.
- How to Read the Siri AI Hands-on: Beta Capabilities, App Intents, and Unsettled Boundaries
Cross-checking The Verge's hands-on with Apple documentation: what iOS 27 Siri AI exposes in developer testing, what third-party apps can prepare, and what still requires beta evidence.
- Google Cloud × CyberLink: Product Decisions for Generative Multimedia AI
A source-checked reading of CyberLink's multimedia AI work: what the Google Cloud case study supports and what product teams still need to validate.
- GPT-Live Voice Architecture: Full-Duplex Interaction, Delegation, and API Boundaries
A grounded look at GPT-Live's full-duplex and background-delegation design in ChatGPT Voice, how it differs from the Realtime API, and which failure modes voice teams should test.
- Meta Muse Spark Through 1.2: Multimodal Reasoning, Parallel Agents, and Version Boundaries
From the original Muse Spark through 1.1 and 1.2, this article separates Meta's published multimodal, parallel-agent, coding, and API capabilities from undocumented internals.
- What Is Meta Muse Image? Agentic Generation, Availability, and Limits
A source-backed analysis of Muse Image's search, coding tools, self-refinement, and test-time compute, with its product availability, evidence, and adoption limits.
- Nano Banana 2 Lite and Gemini Omni Flash: Official Specs, Preview Limits, and Adoption Checks
A grounded review of Google's published specs, pricing, and preview limitations for Nano Banana 2 Lite and Gemini Omni Flash, plus what to validate before adopting an image-to-video pipeline.
- What Is PixelRAG: Retrieval Over Webpage Screenshots
PixelRAG shifts retrieval from plain text to webpage screenshot pixels. This post covers what it solves and where the evidence stops—not a "screenshots beat text" slogan.
- BloomRender User Manual: From Text-to-Image to ID Photos, Portraits, Travel Photos, and Virtual Try-Ons
BloomRender is an AI photo studio powered by Google Gemini. This article details the complete workflow and recommended learning path for text-to-image generation, AI ID photos, editor fine-tuning, portraits, travel photos, and virtual try-ons.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact