init commit
This commit is contained in:
20
docs/CLAUDE.md
Normal file
20
docs/CLAUDE.md
Normal file
@@ -0,0 +1,20 @@
|
||||
# Documentation Folder
|
||||
|
||||
This folder contains business planning, architecture decisions, and documentation.
|
||||
|
||||
**No application code belongs here.**
|
||||
|
||||
## Contents
|
||||
|
||||
| Document | Purpose |
|
||||
|----------|---------|
|
||||
| `code_guidelinees.md` | Code creation guidelines and Git strategy |
|
||||
| `security.md` | Python focused security code suggestions |
|
||||
|
||||
|
||||
## Guidelines
|
||||
|
||||
- Keep documents focused and concise
|
||||
- Update docs when architecture decisions change
|
||||
- Use markdown tables for structured information
|
||||
- Include links to external documentation where relevant
|
||||
144
docs/ROADMAP.md
Normal file
144
docs/ROADMAP.md
Normal file
@@ -0,0 +1,144 @@
|
||||
# SneakyCode Implementation Roadmap
|
||||
|
||||
A phased plan progressing from bare-bones foundation to full autonomous coding agent.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Foundation: Models, Config, and Utilities
|
||||
|
||||
Establish the data layer and shared infrastructure everything else builds on.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/models/config.py` | Pydantic v2 config model — load and validate `config/config.yaml` |
|
||||
| `app/models/message.py` | Message schema (role, content, tool_calls) |
|
||||
| `app/models/tool_call.py` | ToolCall and ToolResult schemas |
|
||||
| `app/utils/logging.py` | Centralized logger with Rich handler |
|
||||
| `app/utils/display.py` | Rich console output helpers (stub — expanded in Phase 2) |
|
||||
| `app/utils/file_helpers.py` | Safe path resolution, binary detection, size guards |
|
||||
| `app/utils/token_counter.py` | Approximate token usage tracking (character-based heuristic for v1) |
|
||||
| `app/main.py` | Entrypoint stub — arg parsing, config load, Rich console setup |
|
||||
|
||||
**Exit criteria:** `python -m app.main --help` runs, config loads and validates, models can be instantiated and serialized.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — TUI and Interactive Shell
|
||||
|
||||
Get a working interactive terminal before wiring up the LLM.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/main.py` | Rich-based interactive REPL loop — prompt for user input, display responses |
|
||||
| `app/utils/display.py` | Formatted output for agent messages, tool calls, errors, token usage |
|
||||
| `app/agent/context.py` | Session state and conversation history management |
|
||||
|
||||
**Exit criteria:** User can type messages into a styled REPL, see them echoed back with formatting, and conversation history is tracked in memory.
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — LLM Integration (Ollama)
|
||||
|
||||
Connect to the local LLM and stream responses into the TUI.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/services/llm.py` | Async httpx client wrapping Ollama's OpenAI-compatible `/v1/chat/completions` endpoint |
|
||||
| `app/services/streaming.py` | SSE parsing, Rich live display, tool call extraction from accumulated stream |
|
||||
|
||||
**Integration:** Wire LLM into the REPL — user message goes to LLM, streamed response displays in real time.
|
||||
|
||||
**Exit criteria:** User can chat with the local model through the TUI with streamed output. Tool call JSON is parsed from the stream but not yet executed.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — Tool Framework and Core Tools
|
||||
|
||||
Build the tool abstraction and implement safe, read-only tools first.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/tools/base.py` | `BaseTool` ABC and `ToolResult` dataclass |
|
||||
| `app/tools/registry.py` | Tool registration, discovery, and JSON schema export for LLM system prompt |
|
||||
| `app/services/permissions.py` | Two-tier approval gating (auto-approve reads; prompt for writes/deletes/shell) |
|
||||
| `app/tools/filesystem.py` | `read_file`, `list_dir` |
|
||||
| `app/tools/search.py` | `grep_files`, `find_files` |
|
||||
|
||||
**Exit criteria:** Tools register themselves, schemas export correctly for inclusion in the system prompt, read-only tools execute and return `ToolResult` objects. Permissions service gates execution.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — Agent Loop (ReAct)
|
||||
|
||||
The core autonomy layer — reason, act, observe, repeat.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/agent/loop.py` | ReAct cycle: send conversation to LLM, parse tool calls, execute, feed results back, repeat |
|
||||
|
||||
**Key behaviors:**
|
||||
- System prompt constructed with tool schemas from registry
|
||||
- Permissions checks before each tool execution
|
||||
- Loop termination on: plain-text response (no tool calls), explicit `finish` tool call, or `max_iterations` exceeded
|
||||
|
||||
**Exit criteria:** Agent can autonomously answer questions about the codebase by chaining `read_file`, `list_dir`, `grep_files`, and `find_files` tool calls in a multi-turn loop.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6 — Write Tools and Shell
|
||||
|
||||
Unlock the agent's ability to modify code and run commands.
|
||||
|
||||
| File | Description |
|
||||
|------|-------------|
|
||||
| `app/tools/filesystem.py` | `write_file`, `make_dir`, `delete_file` (additions to existing module) |
|
||||
| `app/tools/edit.py` | `str_replace` (unique-match required), `patch_apply` |
|
||||
| `app/tools/shell.py` | `run_command` with command allow/deny lists and output truncation |
|
||||
|
||||
**All write/shell operations gated through permissions service.**
|
||||
|
||||
**Exit criteria:** Agent can autonomously create files, edit code via string replacement, and run shell commands — all with user approval for destructive operations.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7 — Polish and Hardening
|
||||
|
||||
Production-readiness: error handling, resource limits, and documentation.
|
||||
|
||||
| Area | Description |
|
||||
|------|-------------|
|
||||
| Error handling | Recovery from malformed tool calls, LLM errors, network timeouts in agent loop |
|
||||
| Token budget | Conversation truncation or summarization when approaching context limit |
|
||||
| Graceful shutdown | Clean Ctrl+C handling, session state preservation |
|
||||
| Testing | End-to-end integration tests (`tests/integration/`), unit tests (`tests/unit/`) |
|
||||
| Documentation | `README.md` with setup and usage instructions, `docs/tools.md` tool reference |
|
||||
|
||||
**Exit criteria:** Agent handles edge cases gracefully, tests pass, and a new user can set up and use the project from the README alone.
|
||||
|
||||
---
|
||||
|
||||
## File Coverage
|
||||
|
||||
Every file from the project structure in CLAUDE.md is accounted for:
|
||||
|
||||
| File | Phase |
|
||||
|------|-------|
|
||||
| `app/main.py` | 1, 2 |
|
||||
| `app/models/config.py` | 1 |
|
||||
| `app/models/message.py` | 1 |
|
||||
| `app/models/tool_call.py` | 1 |
|
||||
| `app/utils/logging.py` | 1 |
|
||||
| `app/utils/display.py` | 1, 2 |
|
||||
| `app/utils/file_helpers.py` | 1 |
|
||||
| `app/utils/token_counter.py` | 1 |
|
||||
| `app/agent/context.py` | 2 |
|
||||
| `app/services/llm.py` | 3 |
|
||||
| `app/services/streaming.py` | 3 |
|
||||
| `app/tools/base.py` | 4 |
|
||||
| `app/tools/registry.py` | 4 |
|
||||
| `app/services/permissions.py` | 4 |
|
||||
| `app/tools/filesystem.py` | 4, 6 |
|
||||
| `app/tools/search.py` | 4 |
|
||||
| `app/agent/loop.py` | 5 |
|
||||
| `app/tools/edit.py` | 6 |
|
||||
| `app/tools/shell.py` | 6 |
|
||||
120
docs/code_guidelines.md
Normal file
120
docs/code_guidelines.md
Normal file
@@ -0,0 +1,120 @@
|
||||
### Coding Standards
|
||||
|
||||
**Style & Structure**
|
||||
- Prefer longer, explicit code over compact one-liners
|
||||
- Always include docstrings for functions/classes + inline comments
|
||||
- Strongly prefer OOP-style code (classes over functional/nested functions)
|
||||
- Strong typing throughout (dataclasses, TypedDict, Enums, type hints)
|
||||
- Value future-proofing and expanded usage insights
|
||||
|
||||
**Data Design**
|
||||
- Use dataclasses for internal data modeling
|
||||
- Typed JSON structures
|
||||
- Functions return fully typed objects (no loose dicts)
|
||||
- Snapshot files in JSON or YAML
|
||||
- Human-readable fields (e.g., `scan_duration`)
|
||||
|
||||
**Templates & UI**
|
||||
- Don't mix large HTML/CSS blocks in Python code
|
||||
- Prefer Jinja templates for HTML rendering
|
||||
- Clean CSS, minimal inline clutter, readable template logic
|
||||
|
||||
**Writing & Documentation**
|
||||
- Markdown documentation
|
||||
- Clear section headers
|
||||
- Roadmap/Phase/Feature-Session style documents
|
||||
- Boilerplate templates first, then refinements
|
||||
|
||||
**Logging**
|
||||
- Use structlog (pip package)
|
||||
- Setup logging at app start: `logger = logging.get_logger(__file__)`
|
||||
|
||||
**Preferred Pip Packages**
|
||||
- API/Web Server: Flask
|
||||
- HTTP: Requests
|
||||
- Logging: Structlog
|
||||
- Scheduling: APScheduler
|
||||
|
||||
### Error Handling
|
||||
- Custom exception classes for domain-specific errors
|
||||
- Consistent error response formats (JSON structure)
|
||||
- Logging severity levels (ERROR vs WARNING)
|
||||
|
||||
### Configuration
|
||||
- Each component has environment-specific configs in its own `/config/*.yaml`
|
||||
- API: `/api/config/development.yaml`, `/api/config/production.yaml`
|
||||
- Web: `/public_web/config/development.yaml`, `/public_web/config/production.yaml`
|
||||
- `.env` for secrets (never committed)
|
||||
- Maintain `.env.example` in each component for documentation
|
||||
- Typed config loaders using dataclasses
|
||||
- Validation on startup
|
||||
|
||||
### Containerization & Deployment
|
||||
- Explicit Dockerfiles
|
||||
- Production-friendly hardening (distroless/slim when meaningful)
|
||||
- Clear build/push scripts that:
|
||||
- Use git branch as tag
|
||||
- Ask whether to tag `:latest`
|
||||
- Ask whether to push
|
||||
- Support private registries
|
||||
|
||||
### API Design
|
||||
- RESTful conventions
|
||||
- Versioning strategy (`/api/v1/...`)
|
||||
- Standardized response format:
|
||||
|
||||
```json
|
||||
{
|
||||
"app": "<APP NAME>",
|
||||
"version": "<APP VERSION>",
|
||||
"status": <HTTP STATUS CODE>,
|
||||
"timestamp": "<UTC ISO8601>",
|
||||
"request_id": "<optional request id>",
|
||||
"result": <data OR null>,
|
||||
"error": {
|
||||
"code": "<optional machine code>",
|
||||
"message": "<human message>",
|
||||
"details": {}
|
||||
},
|
||||
"meta": {}
|
||||
}
|
||||
```
|
||||
|
||||
### Dependency Management
|
||||
- Use `requirements.txt` and virtual environments (`python3 -m venv venv`)
|
||||
- Use path `venv` for all virtual environments
|
||||
- Pin versions to version ranges
|
||||
- Activate venv before running code (unless in Docker)
|
||||
|
||||
### Testing Standards
|
||||
- Manual testing preferred for applications
|
||||
- **API Backend:** Maintain `api/docs/API_TESTING.md` with endpoint examples, curl/httpie commands, expected responses
|
||||
- **Unit tests:** Use pytest for API backend (`api/tests/`)
|
||||
- **Web Frontend:** If using a web frontend, Manual testing checklist are created in `public_web/docs`
|
||||
|
||||
### Git Standards
|
||||
|
||||
**Branch Strategy:**
|
||||
- `master` - Production-ready code only
|
||||
- `dev` - Main development branch, integration point
|
||||
- `beta` - (Optional) Public pre-release testing
|
||||
|
||||
**Workflow:**
|
||||
- Feature work branches off `dev` (e.g., `feature/add-scheduler`)
|
||||
- Merge features back to `dev` for testing
|
||||
- Promote `dev` → `beta` for public testing (when applicable)
|
||||
- Promote `beta` (or `dev`) → `master` for production
|
||||
|
||||
**Commit Messages:**
|
||||
- Use conventional commit format: `feat:`, `fix:`, `docs:`, `refactor:`, etc.
|
||||
- Keep commits atomic and focused
|
||||
- Write clear, descriptive messages
|
||||
|
||||
**Tagging:**
|
||||
- Tag releases on `master` with semantic versioning (e.g., `v1.2.3`)
|
||||
- Optionally tag beta releases (e.g., `v1.2.3-beta.1`)
|
||||
|
||||
---
|
||||
|
||||
## Workflow Preference
|
||||
I follow a pattern: **brainstorm → design → code → revise**
|
||||
46
docs/security.md
Normal file
46
docs/security.md
Normal file
@@ -0,0 +1,46 @@
|
||||
## Foundational Security Instructions
|
||||
|
||||
- Act as a security-aware software engineer generating secure Python code.
|
||||
- Produce implementations that are **secure-by-design and secure-by-default**, not merely cosmetically "secured."
|
||||
- Focus on **preventing vulnerabilities**, not renaming functions or adding superficial security wrappers.
|
||||
- Explicitly identify **trust boundaries** (user input, external systems, internal components) and apply stricter controls at all boundary crossings.
|
||||
- Treat **all external input as untrusted by default**, regardless of source, and validate or sanitize it before use.
|
||||
- Explicitly consider **data sensitivity** (e.g., public, internal, confidential, regulated) and enforce controls appropriate to the highest sensitivity level involved.
|
||||
- Clearly distinguish between **authentication**, **authorization**, and **session management**, and never conflate their responsibilities.
|
||||
- Ensure implementations **fail securely**: errors, exceptions, and edge cases MUST NOT expose sensitive data or weaken security guarantees.
|
||||
- Use inline comments (when generating code) to clearly highlight critical security controls, assumptions, and security-relevant design decisions.
|
||||
- Adhere strictly to OWASP best practices, with particular consideration for the OWASP ASVS.
|
||||
- **Avoid slopsquatting and dependency confusion**: never guess package names or APIs; only reference well-known, reputable, and maintained libraries. Explicitly note any uncommon or low-reputation dependencies.
|
||||
- Do not hardcode secrets, credentials, tokens, or cryptographic material. Always require secure external configuration or secret management mechanisms.
|
||||
|
||||
---
|
||||
|
||||
## Common Weaknesses for Python
|
||||
|
||||
### CWE-79: Improper Neutralization of Input During Web Page Generation ('Cross-site Scripting')
|
||||
**Summary:** Failure to properly sanitize or encode user input can lead to injection of malicious scripts into web pages, enabling XSS attacks.
|
||||
**Mitigation Rule:** All user input rendered in web pages MUST be sanitized and contextually encoded using a secure library such as `bleach` or `html.escape`.
|
||||
|
||||
### CWE-89: Improper Neutralization of Special Elements used in an SQL Command ('SQL Injection')
|
||||
**Summary:** Unsanitized user input in SQL queries can allow attackers to execute arbitrary SQL commands, compromising data integrity and confidentiality.
|
||||
**Mitigation Rule:** SQL queries MUST use parameterized statements or prepared statements provided by libraries such as `sqlite3` or `SQLAlchemy`. Direct concatenation of user input into queries MUST NOT be used.
|
||||
|
||||
### CWE-327: Use of a Broken or Risky Cryptographic Algorithm
|
||||
**Summary:** Using outdated or insecure cryptographic algorithms can compromise data confidentiality and integrity.
|
||||
**Mitigation Rule:** Cryptographic operations MUST use secure algorithms provided by the `cryptography` library. Deprecated algorithms such as MD5 or SHA-1 MUST NOT be used.
|
||||
|
||||
### CWE-798: Use of Hard-coded Credentials
|
||||
**Summary:** Hardcoding credentials in source code can lead to unauthorized access if the code is exposed or leaked.
|
||||
**Mitigation Rule:** Secrets, credentials, and tokens MUST be stored securely using environment variables, secret management tools, or configuration files outside the source code repository.
|
||||
|
||||
### CWE-200: Exposure of Sensitive Information to an Unauthorized Actor
|
||||
**Summary:** Improper error handling or logging can expose sensitive data to unauthorized users.
|
||||
**Mitigation Rule:** Error messages and logs MUST NOT include sensitive information such as stack traces, database connection strings, or user credentials. Use logging libraries such as `logging` with appropriate log levels and sanitization.
|
||||
|
||||
### CWE-502: Deserialization of Untrusted Data
|
||||
**Summary:** Deserializing untrusted data can lead to arbitrary code execution or data tampering.
|
||||
**Mitigation Rule:** Deserialization MUST only be performed on trusted data sources. Unsafe libraries such as `pickle` MUST NOT be used for deserialization of untrusted input.
|
||||
|
||||
### CWE-829: Inclusion of Functionality from Untrusted Control Sphere
|
||||
**Summary:** Using dependencies or code from untrusted sources can introduce malicious functionality or vulnerabilities.
|
||||
**Mitigation Rule:** Dependencies MUST be sourced from reputable package repositories such as PyPI. Verify the integrity and reputation of packages before use, and pin dependency versions to avoid supply chain attacks.
|
||||
Reference in New Issue
Block a user