Harnessing your coding agent to become omni-native.
Omni-IO is a plug-and-play skill set and MCP runtime for understanding, creating, and coordinating images, video, audio, documents, 3D assets, code, and web content—all from natural-language requests.
News · Capabilities · Examples · Repository Structure · Quick Start · How It Works · Citation
Omni-IO-Skills_pr_video_readme.mp4
- [Sep 2026] We release our technical report, Omni-IO Skills: Harnessing Your Agent Omni-Native.
- [Aug 2026] We release Omni-IO Skills, a plug-and-play agent harness for omni-modal understanding, generation, and cross-modal workflows.
Omni is about the breadth of work your agent can accomplish. Start with a real goal—not a tool or file format—and Omni-IO can understand the inputs, plan the work, create the required assets, assemble complete deliverables, inspect the result, and reuse it in whatever comes next.
| Scenario | Tasks you can hand to Omni-IO |
|---|---|
| Office & meetings | Transcribe meeting or interview recordings; organize minutes, decisions, and action items; turn source materials into Word or PDF reports; build company and training presentations; structure tabular data as Excel workbooks. |
| Research & knowledge work | Search and browse the web; extract content from PDFs, Word files, presentations, and spreadsheets; OCR images; analyze source material; synthesize findings into summaries, briefs, or reports. |
| Marketing & social media | Plan platform-specific content; write campaign and post copy; create product images, social covers, and promotional posters; produce short videos with narration, music, and captions; assemble multi-asset campaign packages. |
| Visual design & communication | Generate illustrations and concept art; design invitations, posters, covers, and promotional graphics with accurate text layout; inspect finished designs and selectively rework weak visual or typography layers. |
| Video & audio production | Analyze video scenes, subtitles, speech, and audio content; write scripts and storyboards; generate video shots, narration, music, and sound effects; assemble, caption, inspect, and refine complete videos. |
| Education & training | Turn a topic, textbook, outline, or existing material into course structures, training decks, handouts, explanatory diagrams, narration, and complete educational or popular-science videos. |
| Career & personal branding | Restructure an existing resume; create targeted resumes and cover letters; build interview or self-introduction presentations; produce spoken introductions and complete personal-brand videos. |
| Events & celebrations | Create invitations and event posters with locked dates, names, and venues; generate ambient music and atmosphere visuals; turn event photos into teaser or recap videos. |
| Games, 3D & interactive experiences | Develop character, prop, or environment concept art; turn a reference into a GLB 3D model; create rotation clips or multi-shot showcases; combine assets in an interactive presentation webpage. |
| Custom multi-step projects | Combine any of these tasks into a dependency-aware workflow; run independent work in parallel; pass outputs into downstream tasks; inspect professional deliverables; reuse registered assets across later requests and sessions. |
These scenarios are examples, not a closed list of templates. A request can use one task, combine several functions across scenarios, or become a complete multi-deliverable workflow.
Modalities are the building blocks rather than the starting point. Any of them can serve as an input, an output, or an intermediate asset inside a larger task.
| Modality | Understand | Create |
|---|---|---|
| Text, documents & code | Extract and organize PDF, Word, PowerPoint, Excel, Markdown, code, and webpage content | Generate Markdown, Word, PDF, PowerPoint, Excel, code, and webpages |
| Image | Visual analysis, description, and OCR | Text-to-image generation |
| Video | Scene analysis and subtitle extraction | Video clips and complete multi-shot productions |
| Audio | Speech recognition and audio classification | Speech, music, and sound effects |
| 3D | Geometry inspection and rendered previews | GLB asset generation |
Twenty cases pair source media with finished artifacts. Each card links to its prompt, inputs, outputs, and screen recording.
![]() 01 · Product launch Prompt · Inputs · Outputs · Video |
![]() 03 · Fossil exhibit Prompt · Inputs · Outputs · Video |
![]() 07 · Glowing garden · two turns Prompt · Inputs · Outputs · Video |
![]() 11 · Viking artifact exhibit Prompt · Inputs · Outputs · Video |
![]() 13 · Truffle ingredient story Prompt · Inputs · Outputs · Video |
![]() 18 · Cinematic trailer · two turns Prompt · Inputs · Outputs · Video |
![]() 02 · Illustration masterclass Prompt · Inputs · Outputs · Video |
![]() 08 · Bird field lesson · two turns Prompt · Inputs · Outputs · Video |
![]() 09 · Wallace Line explainer Prompt · Inputs · Outputs · Video |
![]() 10 · Pet care micro-course Prompt · Inputs · Outputs · Video |
![]() 15 · Climate anomaly explainer Prompt · Inputs · Outputs · Video |
![]() 16 · Zealandia planetarium · two turns Prompt · Inputs · Outputs · Video |
![]() 17 · Cell biology gallery · two turns Prompt · Inputs · Outputs · Video |
![]() 19 · Force and motion exhibit Prompt · Inputs · Outputs · Video |
![]() 04 · Avocado 3D asset Prompt · Inputs · Outputs · Video |
![]() 05 · Scan restoration Prompt · Inputs · Outputs · Video |
![]() 14 · Molecular 3D explorer Prompt · Inputs · Outputs · Video |
![]() 20 · Food scan 3D viewer Prompt · Inputs · Outputs · Video |
![]() 06 · Volleyball recap · two turns Prompt · Inputs · Outputs · Video |
![]() 12 · Chess comeback Prompt · Inputs · Outputs · Video |
Omni-IO-Skills/
├── SKILL.md # Entry point and routing index
├── skill.yaml # Skill and MCP declaration
├── skills/ # Atomic understanding, generation, and utility skills
├── expert/ # Professional single-deliverable workflows
├── scenarios/ # Multi-deliverable application workflows
├── orchestration/ # Dependencies, review, interruption, and asset reuse
├── mcp/ # Multimodal tool server
├── config/ # Provider bindings and local credentials
├── formats/ # Internal task declaration formats
├── setup/ # Setup and provider documentation
└── assets/ # README visuals and media
Omni-IO works with hosts that can load an Agent Skills directory and connect to a local stdio MCP server. The commands below support macOS, Linux, and WSL and require Python 3.10+.
Tip
Open this repository in Codex or Claude Code and ask: “Install Omni-IO Skills for this host by following the README Quick Start.” The agent can perform the local setup and tell you which optional credentials are needed for the capabilities you choose.
git clone https://github.com/any2any-mllm/Omni-IO-Skill.git Omni-IO-Skills
cd Omni-IO-Skills
python -m venv .venv
source .venv/bin/activate
python -m pip install -r mcp/requirements.txt
playwright install chromiumplaywright install chromium is needed only for the local browsing capability.
cp config/.env.example config/.envChoose a persistent output directory with OMNI_OUTPUT_DIR. If you do not set one, Omni-IO uses ~/Documents/OmniIO.
You do not need every API key. Start with the capabilities you want, add only their credentials to config/.env, and select providers in config/config.yaml. See the feature list and provider guide for the available options.
Important
Keep config/.env private. It is ignored by Git and should never be committed.
Codex
mkdir -p "$HOME/.agents/skills"
ln -s "$(pwd)" "$HOME/.agents/skills/omni-io"
codex mcp add omni-io \
--env OMNI_OUTPUT_DIR="$HOME/Documents/OmniIO" \
-- "$(pwd)/.venv/bin/python" "$(pwd)/mcp/server.py"
codex mcp listStart a new Codex session after setup.
Claude Code
mkdir -p "$HOME/.claude/skills"
ln -s "$(pwd)" "$HOME/.claude/skills/omni-io"
claude mcp add omni-io \
--scope user \
-e OMNI_OUTPUT_DIR="$HOME/Documents/OmniIO" \
-- "$(pwd)/.venv/bin/python" "$(pwd)/mcp/server.py"
claude mcp listQuit and reopen Claude Code after setup.
Other Agent Skills + MCP hosts
Point the host's skill loader to this repository's root SKILL.md, then register a local stdio MCP server with:
| Field | Value |
|---|---|
| Command | /absolute/path/to/Omni-IO-Skills/.venv/bin/python |
| Arguments | /absolute/path/to/Omni-IO-Skills/mcp/server.py |
| Environment | OMNI_OUTPUT_DIR=/absolute/path/to/a/persistent/workspace |
The host must expose both the Skill and the MCP server to the same agent session.
Start a new conversation and send:
Check my Omni-IO configuration and tell me which capabilities are ready.
A working setup reports which capabilities are ready, which need credentials, and which have key-free paths. If the MCP server connects but the Skill does not appear, verify the symlink and restart the host.
For a longer walkthrough, see setup/quickstart.md.
Most agent setups treat every media tool as a separate integration. Omni-IO gives the host agent one coherent workflow: it understands the request, selects the right skills and providers, executes independent work in parallel, reviews professional deliverables, and keeps successful outputs available for later reuse.
| Understand | Create | Coordinate | Reuse |
|---|---|---|---|
| Inspect images, video, audio, documents, and 3D assets | Generate media, documents, code, and complete deliverables | Plan multi-step jobs with dependency-aware execution | Refer to prior outputs naturally across tasks and restarts |
Omni-IO keeps your existing agent as the reasoning core. The repository adds the skills, tools, provider routing, and persistent asset layer around it.
Omni-IO routes each request through the level of coordination it actually needs: Scenario Skills organize broad multi-deliverable goals, Expert Skills professionally assemble and review a finished artifact, and Atomic Skills perform focused understanding, generation, and utility operations.
- The Skill interprets the request. It selects an atomic, expert, or scenario workflow and builds the task dependencies.
- The MCP runtime executes the work. Understanding, generation, and utility tools expose a consistent interface to the host agent.
- Configuration selects providers. Each capability can use the provider configured in
config/config.yamlwithout changing the higher-level workflow. - The asset registry preserves outputs. Successful files receive stable records so later requests can find and reuse them.
Omni-IO represents multi-step work as an internal dependency graph. Independent tasks can run in parallel; dependent tasks wait for their inputs. Expert workflows inspect the finished deliverable and selectively redo only the parts that failed review.
Generated files are stored under OMNI_OUTPUT_DIR together with registry.json. You can refer naturally to “the last image,” “the previous video,” or another registered output without generating it again. Reuse continues across conversations and host restarts as long as the same output directory is used.
Omni-IO separates the workflow from the provider. You can change a backend in config/config.yaml while keeping the same user-facing request and skill behavior.
- Some capabilities can run locally or through host-native features, including document generation, code and Markdown generation, Edge TTS speech, local 3D inspection, and Playwright browsing.
- Model-backed generation, media understanding, transcription, and search can be enabled individually with provider credentials.
- Provider availability, pricing, and free allowances can change. Review
setup/feature_list.mdandsetup/api_guide.mdbefore choosing a backend. - Changes to
config/.envorconfig/config.yamltake effect after the MCP process restarts.
To isolate separate projects, assign each one a different OMNI_OUTPUT_DIR.
Do I need every API key?
No. Configure only the capabilities you plan to use. The initial configuration check identifies what is ready and what needs a credential, without blocking unrelated capabilities.
Why did the Skill not trigger automatically?
Confirm that the omni-io symlink points to this repository, then restart the host. The request must also involve multimedia understanding, generation, transformation, or a related multi-deliverable workflow.
What should I do if the MCP server is disconnected?
Run the server directly to expose the startup error:
source .venv/bin/activate
python mcp/server.pyMissing packages usually mean the virtual environment is inactive or mcp/requirements.txt was not installed. Fix configuration errors in config/.env or config/config.yaml, then restart the host.
Why did changes to keys or providers not take effect?
The MCP process reads config/.env and config/config.yaml when it starts. Restart the host—or just the omni-io MCP server—after changing either file.
Does Omni-IO replace my model or agent?
No. Your host agent remains the reasoning core. Omni-IO adds reusable instructions, multimodal tools, provider selection, workflow coordination, and persistent artifacts around it.
Can I use only one capability?
Yes. A one-step request routes directly to the relevant Atomic Skill. Expert and Scenario workflows activate only when the requested outcome needs broader production or coordination.
Where are generated files stored?
Under OMNI_OUTPUT_DIR, together with the persistent registry.json. The default is ~/Documents/OmniIO.
Questions, bug reports, feature requests, and workflow ideas are welcome in GitHub Issues. If Omni-IO is useful to you, consider starring the repository so more agent builders can find it.
If you are interested in Omni-IO Skills, please contact Yanlin Li or Hao Fei. If you use this project in your research or applications, please cite our paper:
@article{li2026omniio,
title = {Omni-IO Skills: Harnessing Your Agent Omni-Native},
author = {Li, Yanlin and Hao, Mingyang and Wu, Shengqiong and Fei, Hao and Lee, Mong-Li and Hsu, Wynne},
journal = {arXiv preprint arXiv:2609.31847},
year = {2026}
}






















