- Detect empty LLM responses (no content, no tool calls) instead of
silently treating them as task completion. Retries once without tools
before warning the user.
- Gate /no_think system message and chat_template_kwargs to Qwen/QwQ
models only — sending /no_think to llama3.x caused empty responses.
- Add model_profiles config section for per-model overrides (token
budget, thinking, temperature, max_tokens) matched by name prefix.
Applied at startup and on /model switch.
- Update SessionManager on /model switch so session files record the
correct model.
- Add NDJSON fallback in SSE stream parser for Ollama compatibility.
- Improve read_file error to suggest find_files on FileNotFoundError.
- Add diagnostic logging for empty streams and empty results.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
StreamHandler.reset() was clearing on_content, on_thinking, and on_done
callbacks after every LLM response, but they were only set once per turn.
This caused the thinking indicator and streaming display to stop working
after the first tool call in a multi-step agent turn.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Implement 6 new agent tools — write_file, make_dir, delete_file,
str_replace, patch_apply, run_command — bringing the agent from
read-only observer to active code modifier. All write/shell operations
are gated through the existing permissions service.
Also fix a bug where qwen3.5 thinking mode produces reasoning tokens
but no content after tool results, causing the agent to silently exit.
The loop now detects reasoning-only responses, retries twice, then
injects a nudge message to break the model out of its thinking loop.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>