Skip to content

Latest commit

 

History

16 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Omni-IO Skills logo Omni-IO Skills

Harnessing your coding agent to become omni-native.

Omni-IO is a plug-and-play skill set and MCP runtime for understanding, creating, and coordinating images, video, audio, documents, 3D assets, code, and web content—all from natural-language requests.

GitHub stars Agent Skills compatible MCP stdio server Python 3.10 or newer arXiv paper

News · Capabilities · Examples · Repository Structure · Quick Start · How It Works · Citation


Omni-IO-Skills_pr_video_readme.mp4

News

What You Can Do

Omni is about the breadth of work your agent can accomplish. Start with a real goal—not a tool or file format—and Omni-IO can understand the inputs, plan the work, create the required assets, assemble complete deliverables, inspect the result, and reuse it in whatever comes next.

Omni Tasks

Scenario Tasks you can hand to Omni-IO
Office & meetings Transcribe meeting or interview recordings; organize minutes, decisions, and action items; turn source materials into Word or PDF reports; build company and training presentations; structure tabular data as Excel workbooks.
Research & knowledge work Search and browse the web; extract content from PDFs, Word files, presentations, and spreadsheets; OCR images; analyze source material; synthesize findings into summaries, briefs, or reports.
Marketing & social media Plan platform-specific content; write campaign and post copy; create product images, social covers, and promotional posters; produce short videos with narration, music, and captions; assemble multi-asset campaign packages.
Visual design & communication Generate illustrations and concept art; design invitations, posters, covers, and promotional graphics with accurate text layout; inspect finished designs and selectively rework weak visual or typography layers.
Video & audio production Analyze video scenes, subtitles, speech, and audio content; write scripts and storyboards; generate video shots, narration, music, and sound effects; assemble, caption, inspect, and refine complete videos.
Education & training Turn a topic, textbook, outline, or existing material into course structures, training decks, handouts, explanatory diagrams, narration, and complete educational or popular-science videos.
Career & personal branding Restructure an existing resume; create targeted resumes and cover letters; build interview or self-introduction presentations; produce spoken introductions and complete personal-brand videos.
Events & celebrations Create invitations and event posters with locked dates, names, and venues; generate ambient music and atmosphere visuals; turn event photos into teaser or recap videos.
Games, 3D & interactive experiences Develop character, prop, or environment concept art; turn a reference into a GLB 3D model; create rotation clips or multi-shot showcases; combine assets in an interactive presentation webpage.
Custom multi-step projects Combine any of these tasks into a dependency-aware workflow; run independent work in parallel; pass outputs into downstream tasks; inspect professional deliverables; reuse registered assets across later requests and sessions.

These scenarios are examples, not a closed list of templates. A request can use one task, combine several functions across scenarios, or become a complete multi-deliverable workflow.

Omni Modalities

Modalities are the building blocks rather than the starting point. Any of them can serve as an input, an output, or an intermediate asset inside a larger task.

Modality Understand Create
Text, documents & code Extract and organize PDF, Word, PowerPoint, Excel, Markdown, code, and webpage content Generate Markdown, Word, PDF, PowerPoint, Excel, code, and webpages
Image Visual analysis, description, and OCR Text-to-image generation
Video Scene analysis and subtitle extraction Video clips and complete multi-shot productions
Audio Speech recognition and audio classification Speech, music, and sound effects
3D Geometry inspection and rendered previews GLB asset generation

Examples

Twenty cases pair source media with finished artifacts. Each card links to its prompt, inputs, outputs, and screen recording.

Creative campaigns & visual stories

Product launch: Five bottle photos to Poster, video, and product page
01 · Product launch
Prompt · Inputs · Outputs · Video
Fossil exhibit: Fossil footage and reporting audio to Opening-night campaign
03 · Fossil exhibit
Prompt · Inputs · Outputs · Video
Glowing garden · two turns: Design presentation and narration to Event campaign adapted for a family audience
07 · Glowing garden · two turns
Prompt · Inputs · Outputs · Video
Viking artifact exhibit: Artifact photo and historical narration to Gallery campaign and annotated artifact view
11 · Viking artifact exhibit
Prompt · Inputs · Outputs · Video
Truffle ingredient story: Truffle photo and culinary narration to Recipe story, cards, and storage tips
13 · Truffle ingredient story
Prompt · Inputs · Outputs · Video
Cinematic trailer · two turns: Footage and soundtrack to Trailer recut with a more suspenseful mood
18 · Cinematic trailer · two turns
Prompt · Inputs · Outputs · Video

Learning & explainers

Illustration masterclass: Instructional video and narration to 45-second tutorial and visual teaching materials
02 · Illustration masterclass
Prompt · Inputs · Outputs · Video
Bird field lesson · two turns: Bird photograph and call to Identification lesson adapted for children
08 · Bird field lesson · two turns
Prompt · Inputs · Outputs · Video
Wallace Line explainer: Video and narration to Museum video, maps, comparison, and quiz
09 · Wallace Line explainer
Prompt · Inputs · Outputs · Video
Pet care micro-course: Setup photo, demonstration video, and audio to Training video and care checklist
10 · Pet care micro-course
Prompt · Inputs · Outputs · Video
Climate anomaly explainer: Three sea-temperature anomaly maps to Science-news video and annotated maps
15 · Climate anomaly explainer
Prompt · Inputs · Outputs · Video
Zealandia planetarium · two turns: Geology footage and narration to Planetarium lesson adapted for younger visitors
16 · Zealandia planetarium · two turns
Prompt · Inputs · Outputs · Video
Cell biology gallery · two turns: Animation and narration to Gallery lesson adapted for younger visitors
17 · Cell biology gallery · two turns
Prompt · Inputs · Outputs · Video
Force and motion exhibit: Lab photograph and spoken explanation to Narrated exhibit, experiment card, and quiz
19 · Force and motion exhibit
Prompt · Inputs · Outputs · Video

3D & interactive exhibits

Avocado 3D asset: Spoken food description to Textured model, renders, and interactive viewer
04 · Avocado 3D asset
Prompt · Inputs · Outputs · Video
Scan restoration: Partial point cloud to Reconstructed object and inspection views
05 · Scan restoration
Prompt · Inputs · Outputs · Video
Molecular 3D explorer: Molecular structure image to Labelled model and interactive exhibit
14 · Molecular 3D explorer
Prompt · Inputs · Outputs · Video
Food scan 3D viewer: Textured OBJ scan and reference image to GLB model, inspection views, and viewer
20 · Food scan 3D viewer
Prompt · Inputs · Outputs · Video

Sports & narrative recaps

Volleyball recap · two turns: Match image sequence and commentary to Narrated recap revised for faster viewing
06 · Volleyball recap · two turns
Prompt · Inputs · Outputs · Video
Chess comeback: Fourteen board images to Narrated recap with annotated moves
12 · Chess comeback
Prompt · Inputs · Outputs · Video

Repository Structure

Omni-IO-Skills/
├── SKILL.md                 # Entry point and routing index
├── skill.yaml               # Skill and MCP declaration
├── skills/                  # Atomic understanding, generation, and utility skills
├── expert/                  # Professional single-deliverable workflows
├── scenarios/               # Multi-deliverable application workflows
├── orchestration/           # Dependencies, review, interruption, and asset reuse
├── mcp/                     # Multimodal tool server
├── config/                  # Provider bindings and local credentials
├── formats/                 # Internal task declaration formats
├── setup/                   # Setup and provider documentation
└── assets/                  # README visuals and media

Quick Start

Omni-IO works with hosts that can load an Agent Skills directory and connect to a local stdio MCP server. The commands below support macOS, Linux, and WSL and require Python 3.10+.

Tip

Open this repository in Codex or Claude Code and ask: “Install Omni-IO Skills for this host by following the README Quick Start.” The agent can perform the local setup and tell you which optional credentials are needed for the capabilities you choose.

1. Install the runtime

git clone https://github.com/any2any-mllm/Omni-IO-Skill.git Omni-IO-Skills
cd Omni-IO-Skills

python -m venv .venv
source .venv/bin/activate
python -m pip install -r mcp/requirements.txt
playwright install chromium

playwright install chromium is needed only for the local browsing capability.

2. Create your local configuration

cp config/.env.example config/.env

Choose a persistent output directory with OMNI_OUTPUT_DIR. If you do not set one, Omni-IO uses ~/Documents/OmniIO.

You do not need every API key. Start with the capabilities you want, add only their credentials to config/.env, and select providers in config/config.yaml. See the feature list and provider guide for the available options.

Important

Keep config/.env private. It is ignored by Git and should never be committed.

3. Connect your agent

Codex
mkdir -p "$HOME/.agents/skills"
ln -s "$(pwd)" "$HOME/.agents/skills/omni-io"

codex mcp add omni-io \
  --env OMNI_OUTPUT_DIR="$HOME/Documents/OmniIO" \
  -- "$(pwd)/.venv/bin/python" "$(pwd)/mcp/server.py"

codex mcp list

Start a new Codex session after setup.

Claude Code
mkdir -p "$HOME/.claude/skills"
ln -s "$(pwd)" "$HOME/.claude/skills/omni-io"

claude mcp add omni-io \
  --scope user \
  -e OMNI_OUTPUT_DIR="$HOME/Documents/OmniIO" \
  -- "$(pwd)/.venv/bin/python" "$(pwd)/mcp/server.py"

claude mcp list

Quit and reopen Claude Code after setup.

Other Agent Skills + MCP hosts

Point the host's skill loader to this repository's root SKILL.md, then register a local stdio MCP server with:

Field Value
Command /absolute/path/to/Omni-IO-Skills/.venv/bin/python
Arguments /absolute/path/to/Omni-IO-Skills/mcp/server.py
Environment OMNI_OUTPUT_DIR=/absolute/path/to/a/persistent/workspace

The host must expose both the Skill and the MCP server to the same agent session.

4. Verify the setup

Start a new conversation and send:

Check my Omni-IO configuration and tell me which capabilities are ready.

A working setup reports which capabilities are ready, which need credentials, and which have key-free paths. If the MCP server connects but the Skill does not appear, verify the symlink and restart the host.

For a longer walkthrough, see setup/quickstart.md.

One request. Many modalities. One reusable workflow.

Most agent setups treat every media tool as a separate integration. Omni-IO gives the host agent one coherent workflow: it understands the request, selects the right skills and providers, executes independent work in parallel, reviews professional deliverables, and keeps successful outputs available for later reuse.

Understand Create Coordinate Reuse
Inspect images, video, audio, documents, and 3D assets Generate media, documents, code, and complete deliverables Plan multi-step jobs with dependency-aware execution Refer to prior outputs naturally across tasks and restarts

Omni-IO keeps your existing agent as the reasoning core. The repository adds the skills, tools, provider routing, and persistent asset layer around it.

How It Works

Skill hierarchy

Omni-IO routes each request through the level of coordination it actually needs: Scenario Skills organize broad multi-deliverable goals, Expert Skills professionally assemble and review a finished artifact, and Atomic Skills perform focused understanding, generation, and utility operations.

Scenario, expert, and atomic skill hierarchy in Omni-IO

From intent to execution

Omni-IO architecture overview

  1. The Skill interprets the request. It selects an atomic, expert, or scenario workflow and builds the task dependencies.
  2. The MCP runtime executes the work. Understanding, generation, and utility tools expose a consistent interface to the host agent.
  3. Configuration selects providers. Each capability can use the provider configured in config/config.yaml without changing the higher-level workflow.
  4. The asset registry preserves outputs. Successful files receive stable records so later requests can find and reuse them.

Dependency-aware execution

Omni-IO plans a dependency graph and executes ready tasks in parallel

Omni-IO represents multi-step work as an internal dependency graph. Independent tasks can run in parallel; dependent tasks wait for their inputs. Expert workflows inspect the finished deliverable and selectively redo only the parts that failed review.

Persistent asset reuse

Generated files are stored under OMNI_OUTPUT_DIR together with registry.json. You can refer naturally to “the last image,” “the previous video,” or another registered output without generating it again. Reuse continues across conversations and host restarts as long as the same output directory is used.

Configuration

Omni-IO separates the workflow from the provider. You can change a backend in config/config.yaml while keeping the same user-facing request and skill behavior.

  • Some capabilities can run locally or through host-native features, including document generation, code and Markdown generation, Edge TTS speech, local 3D inspection, and Playwright browsing.
  • Model-backed generation, media understanding, transcription, and search can be enabled individually with provider credentials.
  • Provider availability, pricing, and free allowances can change. Review setup/feature_list.md and setup/api_guide.md before choosing a backend.
  • Changes to config/.env or config/config.yaml take effect after the MCP process restarts.

To isolate separate projects, assign each one a different OMNI_OUTPUT_DIR.

FAQ

Do I need every API key?

No. Configure only the capabilities you plan to use. The initial configuration check identifies what is ready and what needs a credential, without blocking unrelated capabilities.

Why did the Skill not trigger automatically?

Confirm that the omni-io symlink points to this repository, then restart the host. The request must also involve multimedia understanding, generation, transformation, or a related multi-deliverable workflow.

What should I do if the MCP server is disconnected?

Run the server directly to expose the startup error:

source .venv/bin/activate
python mcp/server.py

Missing packages usually mean the virtual environment is inactive or mcp/requirements.txt was not installed. Fix configuration errors in config/.env or config/config.yaml, then restart the host.

Why did changes to keys or providers not take effect?

The MCP process reads config/.env and config/config.yaml when it starts. Restart the host—or just the omni-io MCP server—after changing either file.

Does Omni-IO replace my model or agent?

No. Your host agent remains the reasoning core. Omni-IO adds reusable instructions, multimodal tools, provider selection, workflow coordination, and persistent artifacts around it.

Can I use only one capability?

Yes. A one-step request routes directly to the relevant Atomic Skill. Expert and Scenario workflows activate only when the requested outcome needs broader production or coordination.

Where are generated files stored?

Under OMNI_OUTPUT_DIR, together with the persistent registry.json. The default is ~/Documents/OmniIO.

Community

Questions, bug reports, feature requests, and workflow ideas are welcome in GitHub Issues. If Omni-IO is useful to you, consider starring the repository so more agent builders can find it.

Citation

If you are interested in Omni-IO Skills, please contact Yanlin Li or Hao Fei. If you use this project in your research or applications, please cite our paper:

@article{li2026omniio,
  title   = {Omni-IO Skills: Harnessing Your Agent Omni-Native},
  author  = {Li, Yanlin and Hao, Mingyang and Wu, Shengqiong and Fei, Hao and Lee, Mong-Li and Hsu, Wynne},
  journal = {arXiv preprint arXiv:2609.31847},
  year    = {2026}
}

Star History

Star History Chart

About

No description, website, or topics provided.

Resources

Stars

22 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages