Projects • Prototypes • Experiments
I build things to answer questions research alone can’t.
Small systems, agents, and prototypes used to test an idea before it becomes a roadmap.
Tools used along the way
Instruments, not qualifications.
Evaluation Harness for a Research Repository Agent
A system for making retrieval quality measurable, reproducible, and debatable.
How it works
1. Model the repository
Research team models the scope, topics, and intent of the repository.
2. AI generates questions
AI generates diverse questions based on the modeling.
3. Assign & distribute
Research team members are assigned 20 questions each from the evaluation set.
4. Answer assigned set
Team members answer their assigned questions.
5. Cross-check & debate
Team reviews each other's answers, flags gaps, and debates where needed.
6. Refine & iterate
Consensus builds. We refine the model and improve over time.
The question
Can repository-agent quality be evaluated with evidence rather than “this looks right”?
What I built
- 10 inquiry types
- Evaluation test set
- Cross-checked answer key
- Debate workflow
What I test
- Retrieval quality
- Guardrail behavior
- Source grounding
- Consistency & clarity
Impact
Creates a repeatable, defensible way to measure agent performance and improve it over time.
View project detailsDecision Intelligence Layer
A system that helps teams start with the right context before making the next decision.
1. Start with a problem
User enters a question or decision need.
2. Recommend sources
System identifies relevant enterprise sources.
3. Human confirms
Researcher confirms relevance and importance.
4. Check permissions
Access control honored by infrastructure.
5. Assemble context
Curated context assembled from approved sources.
6. Decision support
LLM provides analysis, options, and next steps.
Why it matters
The right decision depends on the right information — accessible, governed, and usable.
Agent Suite
Two instructional agents built to explore how AI can deepen understanding, not replace effort.
Tiered comprehension agent
…
Design question
How far can the system push the learner upward before we risk discouragement?
Socratic questioning agent
…
Design question
How long can an agent hold the line before productive difficulty becomes frustration?
Shared premise
Learning sticks when the right level of friction meets the right kind of support.