Ivo Pinto’s Post

Do you use response streaming on your Lambdas? Response streaming now supports 200 MB payloads (10x increase from 20 MB) – a massive change for AI applications. What 20 MB → 200 MB unlocks: Text: ~5M → 50M characters (~20K → 200K typical LLM tokens) PDFs: ~200 → 2,000 pages with images Images: ~20 → 200 high-res processed results Audio: ~3 → 30 minutes of processed/enhanced audio files Why this matters for you RAG responses - Return entire document chunks with metadata in one stream Batch inference - Process multiple inputs and stream all results together Audio processing - Full transcription with timestamps, speaker IDs, confidence scores Basically, eliminate bypasses complex chunking logic for outputs exceeding 20 MB. But remember, Lambda still has a 15-minute execution limit. Streaming is an entirely different experience. See results as they're generated, instead of after everything is processed. #aws #lambda #cloudcomputing

  • Batch vs streaming response side-by-side on terminal

To view or add a comment, sign in

Explore content categories