大功告成
This commit is contained in:
@@ -0,0 +1,278 @@
|
||||
# 3D Digital Human System Technical Documentation
|
||||
|
||||
## 1. Project Summary
|
||||
|
||||
This project implements a real-time conversational digital human system with these core capabilities:
|
||||
|
||||
- Text and voice interaction
|
||||
- LLM response generation (DeepSeek online, with fallback)
|
||||
- TTS synthesis for reply audio
|
||||
- Audio-driven 3D animation control stream
|
||||
- Browser-side 3D rendering (Three.js + VRM/GLTF)
|
||||
- WebRTC audio transport and WebSocket animation transport
|
||||
|
||||
Current workspace root:
|
||||
|
||||
- `/home/xsl/code/product`
|
||||
|
||||
## 2. Runtime Architecture
|
||||
|
||||
### 2.1 High-Level Pipeline
|
||||
|
||||
1. User input enters through `/chat/text` or voice pipeline.
|
||||
2. LLM generates reply text.
|
||||
3. TTS synthesizes reply waveform.
|
||||
4. Avatar service converts audio chunks to animation control frames.
|
||||
5. Backend pushes:
|
||||
- reply audio to WebRTC audio track
|
||||
- animation frames to `/ws/animation`
|
||||
6. Frontend renders 3D model and applies mouth/head/body motions.
|
||||
|
||||
### 2.2 Transport Split
|
||||
|
||||
- WebRTC: audio only
|
||||
- WebSocket (`/ws/animation`): animation control data
|
||||
- SSE (`/events`): system metrics and runtime status
|
||||
|
||||
This split removes server-side video generation and makes rendering native in the browser.
|
||||
|
||||
## 3. Core Modules
|
||||
|
||||
### 3.1 Backend Entry
|
||||
|
||||
- `main.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- FastAPI app and route registration
|
||||
- WebRTC offer/peer management
|
||||
- Animation WebSocket broadcasting
|
||||
- Runtime metrics and `/health`
|
||||
- Session reset and chat endpoints
|
||||
|
||||
### 3.2 Pipeline Orchestration
|
||||
|
||||
- `core/pipeline.py`
|
||||
- `core/state_machine.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Turn-based orchestration (input -> LLM -> TTS -> output)
|
||||
- State transitions (idle, thinking, speaking)
|
||||
- Latency accounting and result metadata
|
||||
|
||||
### 3.3 Services
|
||||
|
||||
- `services/llm.py`
|
||||
- `services/tts.py`
|
||||
- `services/asr.py`
|
||||
- `services/vad.py`
|
||||
- `services/avatar.py`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- `LLMService`: online model calls via OpenAI-compatible client
|
||||
- `TTSService`: speech synthesis
|
||||
- `AvatarService`: audio-to-controls mapping for mouth/viseme/head motion
|
||||
|
||||
### 3.4 Frontend
|
||||
|
||||
- `web/index.html`
|
||||
- `web/app.js`
|
||||
- `web/style.css`
|
||||
|
||||
Responsibilities:
|
||||
|
||||
- Three.js scene setup and camera controls
|
||||
- VRM/GLTF model loading (URL or local file)
|
||||
- Animation buffering and frame application
|
||||
- Idle body pose + talking expressions
|
||||
- WebRTC audio playback
|
||||
|
||||
## 4. Animation Data Contract
|
||||
|
||||
Backend pushes animation chunks shaped like:
|
||||
|
||||
- `type`: `animation_chunk` / `animation_reset` / `animation_state`
|
||||
- `fps`, `duration_ms`, `frame_count`
|
||||
- `frames[]` with:
|
||||
- `seq`
|
||||
- `time_ms`
|
||||
- `controls`
|
||||
|
||||
Typical controls include:
|
||||
|
||||
- `jawOpen`
|
||||
- `viseme_aa`, `viseme_ee`, `viseme_oh`
|
||||
- `mouthPucker`
|
||||
- `headYaw`, `headPitch`, `headRoll`
|
||||
|
||||
Frontend behavior:
|
||||
|
||||
- buffers frames for stable playback
|
||||
- applies VRM expressions when available
|
||||
- falls back to morph targets or bone motion
|
||||
|
||||
## 5. Configuration
|
||||
|
||||
Configuration file:
|
||||
|
||||
- `.env` (root)
|
||||
|
||||
### 5.1 LLM Settings
|
||||
|
||||
Supported API key aliases:
|
||||
|
||||
- `LLM_API_KEY`
|
||||
- `DEEPSEEK_API_KEY`
|
||||
- `OPENAI_API_KEY`
|
||||
|
||||
Recommended DeepSeek settings:
|
||||
|
||||
- `LLM_BASE_URL=https://api.deepseek.com/v1`
|
||||
- `LLM_MODEL=deepseek-chat`
|
||||
|
||||
### 5.2 Server Settings
|
||||
|
||||
- `WEBRTC_HOST` (default `0.0.0.0`)
|
||||
- `WEBRTC_PORT` (default `8018`)
|
||||
- `SSL_CERTFILE`
|
||||
- `SSL_KEYFILE`
|
||||
|
||||
### 5.3 Avatar Settings
|
||||
|
||||
- `AVATAR_DRIVER_MODE=blendshape_stream`
|
||||
- `AVATAR_CONTROL_PROTOCOL=ws`
|
||||
- `AVATAR_BLENDSHAPE_SCHEMA=arkit`
|
||||
|
||||
## 6. Startup and Process Control
|
||||
|
||||
New process script:
|
||||
|
||||
- `scripts/service-8018.sh`
|
||||
|
||||
Usage:
|
||||
|
||||
```bash
|
||||
bash scripts/service-8018.sh start
|
||||
bash scripts/service-8018.sh status
|
||||
bash scripts/service-8018.sh logs
|
||||
bash scripts/service-8018.sh restart
|
||||
bash scripts/service-8018.sh stop
|
||||
```
|
||||
|
||||
Behavior:
|
||||
|
||||
- Loads environment from `.env`
|
||||
- Starts uvicorn from project venv if available
|
||||
- Supports HTTPS cert/key from `.env`
|
||||
- Writes pid and log to `.run/`
|
||||
- Cleans stale python listeners on target port
|
||||
|
||||
Runtime files:
|
||||
|
||||
- `.run/visual-chat-8018.pid`
|
||||
- `.run/visual-chat-8018.log`
|
||||
|
||||
## 7. API and Web Endpoints
|
||||
|
||||
### 7.1 Health and Metadata
|
||||
|
||||
- `GET /health`
|
||||
- `GET /meta`
|
||||
|
||||
### 7.2 Chat and Session
|
||||
|
||||
- `POST /chat/text`
|
||||
- `POST /chat/reset`
|
||||
|
||||
### 7.3 Streaming Channels
|
||||
|
||||
- `GET /events` (SSE)
|
||||
- `GET /ws/subtitles` (WebSocket)
|
||||
- `GET /ws/animation` (WebSocket)
|
||||
|
||||
### 7.4 WebRTC
|
||||
|
||||
- `POST /webrtc/offer`
|
||||
|
||||
## 8. Validation Checklist
|
||||
|
||||
### 8.1 Service Up
|
||||
|
||||
- `bash scripts/service-8018.sh status` should show running
|
||||
- `GET /health` should return valid JSON
|
||||
|
||||
### 8.2 LLM Online
|
||||
|
||||
Health fields should indicate:
|
||||
|
||||
- `llm.ready = true`
|
||||
- `llm.error = null`
|
||||
- `last_llm_source = online` after at least one chat turn
|
||||
|
||||
### 8.3 Frontend Motion and Audio
|
||||
|
||||
- Model loads successfully
|
||||
- Mouth motion follows replies
|
||||
- Body remains in non-T-pose idle stance
|
||||
- Remote audio is audible via WebRTC
|
||||
|
||||
## 9. Troubleshooting Guide
|
||||
|
||||
### 9.1 LLM stays fallback
|
||||
|
||||
Symptoms:
|
||||
|
||||
- `last_llm_source = fallback`
|
||||
- `llm.ready = false`
|
||||
|
||||
Checks:
|
||||
|
||||
1. Confirm `.env` exists in workspace root.
|
||||
2. Confirm key value is set in `LLM_API_KEY` (or alias).
|
||||
3. Restart service after any `.env` update.
|
||||
4. Verify `llm.base_url` and `llm.model` in `/health`.
|
||||
|
||||
### 9.2 Model load errors
|
||||
|
||||
Symptoms:
|
||||
|
||||
- Web UI shows model fetch failures
|
||||
|
||||
Checks:
|
||||
|
||||
1. Use reachable URL or local `.vrm` file.
|
||||
2. Confirm browser can access external CDN.
|
||||
3. Prefer VRM full-body models for body-pose control.
|
||||
|
||||
### 9.3 Mouth not moving
|
||||
|
||||
Checks:
|
||||
|
||||
1. Confirm animation websocket connected.
|
||||
2. Confirm animation frame count increases.
|
||||
3. Confirm TTS is ready and not in silence fallback.
|
||||
|
||||
### 9.4 HTTPS/microphone issues on LAN
|
||||
|
||||
Checks:
|
||||
|
||||
1. Use valid cert paths in `.env`.
|
||||
2. Trust self-signed cert on browser/device.
|
||||
3. Ensure firewall allows target port.
|
||||
|
||||
## 10. Security and Operations Notes
|
||||
|
||||
- Keep `.env` out of source control.
|
||||
- Rotate API keys regularly.
|
||||
- Restrict exposed port by network policy if internet-facing.
|
||||
- Consider reverse proxy and managed TLS for production.
|
||||
|
||||
## 11. Suggested Next Production Steps
|
||||
|
||||
1. Add structured request tracing (request id, turn id).
|
||||
2. Add persistent metrics export (Prometheus/OpenTelemetry).
|
||||
3. Add automatic reconnect and health watchdog.
|
||||
4. Add integration tests for LLM online/fallback branches.
|
||||
5. Add model asset mirror to avoid public CDN instability.
|
||||
Reference in New Issue
Block a user