Ph.D. Candidate, Civil & Environmental Engineering, Lehigh University

Wenyu Chiou

AI-agent infrastructure, and the science of when to trust it

AI evaluation / research engineer with catastrophe-risk domain depth.

I build AI-agent infrastructure — MCP tooling, multi-agent orchestration, and agent harnesses — and I apply large language models as the behavioral engine of agent-based models of complex human–water systems, grounded in survey data I collected myself, with the governance that keeps agent behavior trustworthy.

Open to Summer 2027 internships in AI evaluation, research engineering, and risk analytics.

Wenyu Chiou at the AGU Fall Meeting 2025, beside his research poster.

Shipped AI infrastructure

research-hub — an MCP server and CLI (PyPI: research-hub-pipeline v1.1.1) that makes Zotero, Obsidian, and NotebookLM AI-operable. Listed in Awesome MCP Servers; tested on Windows, macOS, and Linux.

Maintained open-source tool

AI-agent skills & governance

Open-source agent skills — ai-research-skills (a 15-skill research catalog), agent-collab-skills, and codex-delegate — plus WAGF, a governance framework that constrains LLM agents so agent-driven simulations stay trustworthy.

Maintained open-source tools · WAGF in preparation

Community & education

awesome-agentic-ai-zh — a trilingual roadmap for building agentic AI, from LLM basics to production multi-agent systems, with automated checks that keep all three languages in sync.

4.7k+ GitHub stars, 622 forks (July 2026)

Selected engineering

research-hub

Maintained open-source tool

Literature workflows are repetitive and easy to get subtly wrong across tools. research-hub turns Zotero, Obsidian, and NotebookLM into one AI-operable workspace — searching, ingesting, and syncing papers through a single CLI, MCP server, and REST API.

codex-delegate

Maintained open-source tool

Delegating code work to a second AI agent is only economical if the delegate cannot fabricate success. codex-delegate verifies the delegate’s completion from git state rather than from its own report.

awesome-agentic-ai-zh

Maintained open-source tool

Chinese-speaking learners lacked a staged path into agentic AI, and hand-mirrored translations drift. This trilingual curriculum keeps its locales verifiably in sync with CI.

4.7k+ GitHub stars (July 2026)

Selected research

FLOODABM

Research prototype

Flood-adaptation models usually assert household behavior instead of measuring it. FLOODABM grounds 52,141 simulated households in a 937-household survey and validates losses against observed insurance claims.

Cat_framework

Research prototype

A loss number without an inspectable validation chain cannot be trusted. This Hazus-based seismic bridge-loss pipeline reports where the recipe fails, not only where it works.

The full research program

Current focus: evaluation and governance methods for LLM agents — testing whether they reproduce real human decision pathways (grounded in a 937-household survey I designed and fielded) and constraining agent-driven simulations so they stay trustworthy. Research software in preparation for release.

Documents

Contact

Open to Summer 2027 internships in AI evaluation, research engineering, and risk analytics.

Email wec324@lehigh.edu

wec324@lehigh.edu