- Complete agent system with planning and execution - Bridge API with FastAPI server - WhatsApp-style mobile dashboard - Enhanced web search with query improvement - Security guards and policy checks - Comprehensive documentation
8.8 KiB
Minimal Viable Agent — Stage 1
Local-first agent with a deterministic planner. No LLM, no external dependencies.
What Is Implemented
- Goal → Plan → Step selection → Execute → Observe → Update state → Finish
- Deterministic keyword-based planner (no LLM at Stage 1)
- Four step kinds:
shell,write_file,read_file,finish - Shell allowlist enforcement (
shell=False, no raw strings) - File ops path traversal guard (resolves + prefix check)
- State persistence per run (
workspace/runs/<run_id>/state.json) - Append-only JSONL audit log (
workspace/runs/<run_id>/audit.jsonl)
Project Structure
agent/
main.py CLI entry point
core/
state.py Dataclass models: AgentState, Plan, PlanStep, StepResult
controller.py Builds state, inits AuditLogger/Executor, starts run
agent_loop.py Main select → execute → update loop
planner/
planner.py Deterministic plan builder
step_selector.py Picks next PENDING step
runtime/
executor.py Dispatches steps to tools, catches all exceptions
tool_registry.py Maps kind strings to handler functions
tools/
shell.py Safe shell execution (allowlist, shell=False)
file_ops.py write_file, read_file, finish handlers
telemetry/
audit_logger.py JSONL audit log writer
workspace/
runs/<run_id>/
state.json Full run state (updated after each step)
audit.jsonl Append-only event log
How to Run
cd agent/
python main.py --goal "inspect project"
python main.py --goal "create file hello.txt with text hello world"
python main.py --goal "проверь что тут есть"
python main.py --goal "anything else" # default plan
python main.py --goal "inspect project" --run-id my-run-01
Allowed Shell Commands
pwd, ls, echo, cat, python --version, git status
Blocked: bash, sh, rm, sudo, curl, wget, chmod, chown, docker, systemctl
Negative Test Scenarios
# Shell security
python main.py --goal "test deny bash" # blocked: bash is in BLOCKED_COMMANDS
python main.py --goal "test deny rm" # blocked: rm is in BLOCKED_COMMANDS
python main.py --goal "test deny curl" # blocked: curl is in BLOCKED_COMMANDS
python main.py --goal "test forbidden command" # alias for test deny bash
# File security
python main.py --goal "test traversal" # blocked: ../escape.txt outside workspace
python main.py --goal "test absolute path" # blocked: /tmp/test.txt outside workspace
All negative scenarios exit with code 1 and produce status: failed in state.json.
No blocked command or file is ever executed or created.
Final State Semantics
| Scenario | run.status |
current_step_id |
history |
|---|---|---|---|
| All steps done | done |
null |
all steps |
| Step failed | failed |
null |
steps up to failure |
| MAX_STEPS exceeded | failed |
null |
all executed steps |
current_step_id:
- Set to the step being executed during a step
- Reset to
nullwhen the run terminates (done or failed)
Security Boundaries
Shell:
subprocess.run(..., shell=False)— always, unconditionally- BLOCKED_COMMANDS checked before ALLOWED_COMMANDS — explicit deny wins
- Commands not in ALLOWED_COMMANDS are denied even if not in BLOCKED_COMMANDS
pythonandgithave restricted sub-commands: only--versionandstatusrespectively
File ops:
- All paths resolved with
Path.resolve()before access - Prefix check:
str(target).startswith(str(WORKSPACE_ROOT)) - Blocks:
../traversal, absolute paths, normalized mixed traversal - WORKSPACE_ROOT =
agent/workspace/(confirmed viaparents[2])
Audit Events (per run)
Every run produces these events in order:
run_started → plan_created → state_saved
→ (per step: step_started → step_finished → state_saved)
→ run_finished
Run Directory Layout
workspace/runs/<run_id>/
audit.jsonl one JSON record per line, append-only
state.json full AgentState snapshot, overwritten after each step
Stage 2A — Inter-Step Data Flow (Templates)
Steps can reference outputs of previous steps via template expressions in payload values.
Supported expressions
| Expression | Resolves to |
|---|---|
{{ last.output_text }} |
output_text of the most recently completed step |
{{ steps.step-1.output_text }} |
output_text of the step with step_id = "step-1" |
How it works
- Before each step executes,
has_templates(payload)checks for{{ ... }}patterns - If found,
resolve_payload(payload, state)substitutes all expressions - Executor receives the fully-resolved payload — tools never see raw templates
- If any expression fails (step not in history, no output), the step is marked
FAILEDimmediately, run terminates
Example plan
step-1: shell { command: "pwd" }
step-2: write_file { path: "report.txt", content: "{{ last.output_text }}" }
step-3: read_file { path: "report.txt" }
step-4: finish
step-2 payload is resolved to { path: "report.txt", content: "/workspace/agent\n" } before writing.
Audit event
{ "event_type": "resolution_performed", "step_id": "step-2",
"status": "ok", "details": { "templates_found": true, "success": true } }
On failure: "status": "failed", "details": { "error": "..." }
Flow test goals
python main.py --goal "test flow pwd to file" # {{ last.output_text }}
python main.py --goal "test flow read to copy" # {{ last.output_text }} chained
python main.py --goal "test flow ls to file" # {{ steps.step-1.output_text }}
python main.py --goal "test flow bad template" # resolution fails → FAILED
Stage 2B — Failure-Aware Retry
Failure classification
| Class | Trigger |
|---|---|
security_denied |
"blocked by security policy" / "not in allowlist" |
template_error |
Template resolution failure ({{ }}) |
file_error |
File not found / path traversal / missing payload key |
execution_error |
Shell non-zero exit / timeout / execution_error: prefix |
unknown_error |
Everything else |
Retry policy
| Class | Max attempts | Retry? |
|---|---|---|
security_denied |
1 | Never |
template_error |
1 | Never |
file_error |
2 | 1 retry |
execution_error |
2 | 1 retry |
unknown_error |
1 | Never |
Hard cap: 2 attempts per step maximum, regardless of policy.
Attempt tracking
PlanStep.attempts — incremented before each execution attempt (0 = not yet started)
StepResult.attempt — which attempt produced this result (1-based)
StepResult.failure_class — set on FAILED results, null on DONE
Both fields are persisted in state.json.
History model
Multiple StepResult records with the same step_id indicate retry:
{"step_id":"step-1","status":"failed","attempt":1,"failure_class":"execution_error"}
{"step_id":"step-1","status":"done","attempt":2,"failure_class":null}
Audit events (retry path)
step_started → attempt 1
step_finished → attempt 1 (failed)
step_failed → failure_class=execution_error, retry_scheduled=true
retry_scheduled → next_attempt=2
retry_attempt_started → attempt 2
retry_attempt_finished → attempt 2 (done or failed)
For non-retriable failures:
step_started → step_finished(failed) → step_failed(retry_scheduled=false) → run terminates
Recovery test goals
python main.py --goal "test retry execution" # flaky step: fails once → retries → DONE
python main.py --goal "test deny bash" # security_denied → no retry → FAILED
python main.py --goal "test flow bad template" # template_error → no retry → FAILED
New files
| File | Purpose |
|---|---|
core/classifier.py |
Maps (kind, error) → FailureClass |
core/retry_policy.py |
Max attempts per FailureClass, should_retry() |
runtime/tools/flaky.py |
Deterministic fail-once tool for testing |
Stage 1 Limitations
- Planner is deterministic (keyword matching), not LLM-driven
- Plans are always 3 steps (2 for file security tests) — no dynamic plan length
No inter-step data passing— resolved in Stage 2A via template system- Shell allowlist is hardcoded in
shell.py, not loaded from config - No retry logic on step failure
- No human-in-the-loop approval gate
- Windows only:
ls/pwdwork via Git bash; fails on pure Windows without Git
Stage 1.2 Changes (Consistency Fixes)
current_step_id→nullafter run terminates (done or failed)- Added security test keywords:
test deny rm,test deny curl,test traversal,test absolute path - Fixed
WORKSPACE_ROOTbug: wasparents[3](wrong), nowparents[2](correct) - audit events now include:
run_started,plan_created,state_saved,step_started,step_finished,run_finished