Files
product/TECHNICAL_DOCUMENTATION.md
T
2026-03-27 17:10:41 +08:00

279 lines
6.0 KiB
Markdown

# 3D Digital Human System Technical Documentation
## 1. Project Summary
This project implements a real-time conversational digital human system with these core capabilities:
- Text and voice interaction
- LLM response generation (DeepSeek online, with fallback)
- TTS synthesis for reply audio
- Audio-driven 3D animation control stream
- Browser-side 3D rendering (Three.js + VRM/GLTF)
- WebRTC audio transport and WebSocket animation transport
Current workspace root:
- `/home/xsl/code/product`
## 2. Runtime Architecture
### 2.1 High-Level Pipeline
1. User input enters through `/chat/text` or voice pipeline.
2. LLM generates reply text.
3. TTS synthesizes reply waveform.
4. Avatar service converts audio chunks to animation control frames.
5. Backend pushes:
- reply audio to WebRTC audio track
- animation frames to `/ws/animation`
6. Frontend renders 3D model and applies mouth/head/body motions.
### 2.2 Transport Split
- WebRTC: audio only
- WebSocket (`/ws/animation`): animation control data
- SSE (`/events`): system metrics and runtime status
This split removes server-side video generation and makes rendering native in the browser.
## 3. Core Modules
### 3.1 Backend Entry
- `main.py`
Responsibilities:
- FastAPI app and route registration
- WebRTC offer/peer management
- Animation WebSocket broadcasting
- Runtime metrics and `/health`
- Session reset and chat endpoints
### 3.2 Pipeline Orchestration
- `core/pipeline.py`
- `core/state_machine.py`
Responsibilities:
- Turn-based orchestration (input -> LLM -> TTS -> output)
- State transitions (idle, thinking, speaking)
- Latency accounting and result metadata
### 3.3 Services
- `services/llm.py`
- `services/tts.py`
- `services/asr.py`
- `services/vad.py`
- `services/avatar.py`
Responsibilities:
- `LLMService`: online model calls via OpenAI-compatible client
- `TTSService`: speech synthesis
- `AvatarService`: audio-to-controls mapping for mouth/viseme/head motion
### 3.4 Frontend
- `web/index.html`
- `web/app.js`
- `web/style.css`
Responsibilities:
- Three.js scene setup and camera controls
- VRM/GLTF model loading (URL or local file)
- Animation buffering and frame application
- Idle body pose + talking expressions
- WebRTC audio playback
## 4. Animation Data Contract
Backend pushes animation chunks shaped like:
- `type`: `animation_chunk` / `animation_reset` / `animation_state`
- `fps`, `duration_ms`, `frame_count`
- `frames[]` with:
- `seq`
- `time_ms`
- `controls`
Typical controls include:
- `jawOpen`
- `viseme_aa`, `viseme_ee`, `viseme_oh`
- `mouthPucker`
- `headYaw`, `headPitch`, `headRoll`
Frontend behavior:
- buffers frames for stable playback
- applies VRM expressions when available
- falls back to morph targets or bone motion
## 5. Configuration
Configuration file:
- `.env` (root)
### 5.1 LLM Settings
Supported API key aliases:
- `LLM_API_KEY`
- `DEEPSEEK_API_KEY`
- `OPENAI_API_KEY`
Recommended DeepSeek settings:
- `LLM_BASE_URL=https://api.deepseek.com/v1`
- `LLM_MODEL=deepseek-chat`
### 5.2 Server Settings
- `WEBRTC_HOST` (default `0.0.0.0`)
- `WEBRTC_PORT` (default `8018`)
- `SSL_CERTFILE`
- `SSL_KEYFILE`
### 5.3 Avatar Settings
- `AVATAR_DRIVER_MODE=blendshape_stream`
- `AVATAR_CONTROL_PROTOCOL=ws`
- `AVATAR_BLENDSHAPE_SCHEMA=arkit`
## 6. Startup and Process Control
New process script:
- `scripts/service-8018.sh`
Usage:
```bash
bash scripts/service-8018.sh start
bash scripts/service-8018.sh status
bash scripts/service-8018.sh logs
bash scripts/service-8018.sh restart
bash scripts/service-8018.sh stop
```
Behavior:
- Loads environment from `.env`
- Starts uvicorn from project venv if available
- Supports HTTPS cert/key from `.env`
- Writes pid and log to `.run/`
- Cleans stale python listeners on target port
Runtime files:
- `.run/visual-chat-8018.pid`
- `.run/visual-chat-8018.log`
## 7. API and Web Endpoints
### 7.1 Health and Metadata
- `GET /health`
- `GET /meta`
### 7.2 Chat and Session
- `POST /chat/text`
- `POST /chat/reset`
### 7.3 Streaming Channels
- `GET /events` (SSE)
- `GET /ws/subtitles` (WebSocket)
- `GET /ws/animation` (WebSocket)
### 7.4 WebRTC
- `POST /webrtc/offer`
## 8. Validation Checklist
### 8.1 Service Up
- `bash scripts/service-8018.sh status` should show running
- `GET /health` should return valid JSON
### 8.2 LLM Online
Health fields should indicate:
- `llm.ready = true`
- `llm.error = null`
- `last_llm_source = online` after at least one chat turn
### 8.3 Frontend Motion and Audio
- Model loads successfully
- Mouth motion follows replies
- Body remains in non-T-pose idle stance
- Remote audio is audible via WebRTC
## 9. Troubleshooting Guide
### 9.1 LLM stays fallback
Symptoms:
- `last_llm_source = fallback`
- `llm.ready = false`
Checks:
1. Confirm `.env` exists in workspace root.
2. Confirm key value is set in `LLM_API_KEY` (or alias).
3. Restart service after any `.env` update.
4. Verify `llm.base_url` and `llm.model` in `/health`.
### 9.2 Model load errors
Symptoms:
- Web UI shows model fetch failures
Checks:
1. Use reachable URL or local `.vrm` file.
2. Confirm browser can access external CDN.
3. Prefer VRM full-body models for body-pose control.
### 9.3 Mouth not moving
Checks:
1. Confirm animation websocket connected.
2. Confirm animation frame count increases.
3. Confirm TTS is ready and not in silence fallback.
### 9.4 HTTPS/microphone issues on LAN
Checks:
1. Use valid cert paths in `.env`.
2. Trust self-signed cert on browser/device.
3. Ensure firewall allows target port.
## 10. Security and Operations Notes
- Keep `.env` out of source control.
- Rotate API keys regularly.
- Restrict exposed port by network policy if internet-facing.
- Consider reverse proxy and managed TLS for production.
## 11. Suggested Next Production Steps
1. Add structured request tracing (request id, turn id).
2. Add persistent metrics export (Prometheus/OpenTelemetry).
3. Add automatic reconnect and health watchdog.
4. Add integration tests for LLM online/fallback branches.
5. Add model asset mirror to avoid public CDN instability.