mirror of
https://github.com/deepseek-ai/deepseek-harness
synced 2026-08-15 21:04:50 +00:00
fix(llm-pi-ai): let a model declare the request modalities it accepts
A model the installed pi-ai catalog does not describe was reported as text-only with no way to say otherwise, so a vision model added through the custom-provider form was refused at every image admission point. The justification in the source described the DeepSeek chat-completions serializer, which does reject image blocks; the pi-ai request converter and every wire protocol it speaks carry images. Modalities now resolve entry `input` -> installed catalog entry -> route `defaultInput`, the chain the two capacity fallbacks already use, so the route value is a fallback and never narrows a catalog model. Its default stays `[text]`: nothing can interrogate a gateway for its modalities, and over-claiming admits an image the provider rejects mid-turn, after prompt admission has already committed the message.
This commit is contained in:
@@ -2,5 +2,5 @@
|
||||
# side as of the last confirmed-consistent state. Both languages carry equal authority;
|
||||
# after editing either side, bring the other along and re-record with:
|
||||
# pnpm run verify-translation-pairing --write docs/user/guide/providers.md
|
||||
providers.md: a3f94f0cc86401c0f9e5b94cfd823bf9f08e6bfc
|
||||
providers.zh.md: 7d74e0086e62d8a0c2fb39085207125b4b4354e7
|
||||
providers.md: 099f434ec4602aa402239e83c708d81fcadd7732
|
||||
providers.zh.md: 367c90b525ad628b3cd86b2d22045c25064e88a1
|
||||
|
||||
@@ -28,6 +28,57 @@ The Provider ID is permanent because requests, saved sessions, model defaults, a
|
||||
|
||||
Under **Model catalog**, choose **Fetch available models** to query the base URL and credential currently shown in the form. Selecting candidates updates the draft; the provider is not stored until you save. Catalog providers use their installed catalog without a network request.
|
||||
|
||||
### Image input
|
||||
|
||||
A model you enter by hand is treated as text-only until it says otherwise, because nothing can ask an endpoint which modalities it accepts. Attaching an image to such a model is refused before it is sent, naming the model.
|
||||
|
||||
A vision model on a custom provider therefore needs one line. The form has no field for it; add `input` to the model in `$DSH_HOME/settings.yaml`:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
my-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://gateway.example/v1
|
||||
models:
|
||||
- id: legacy-chat
|
||||
- id: vision-preview
|
||||
input: [text, image]
|
||||
```
|
||||
|
||||
`input` accepts `text` and `image`, and applies to that model alone, so one route can serve both kinds. Omitting it — or writing an empty list, which means the same thing — keeps whatever the installed catalog records for that model, and falls back to the route's `defaultInput` for a model the catalog does not describe.
|
||||
|
||||
If every model you entered by hand takes images, set the fallback once on the route instead of on each of them:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
vision-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://vision.example/v1
|
||||
defaultInput: [text, image]
|
||||
models:
|
||||
- id: first-model
|
||||
- id: second-model
|
||||
```
|
||||
|
||||
`defaultInput` is a fallback, not an override, and defaults to `[text]`: on a catalog provider it answers only for models the catalog does not describe, so it never removes images from a catalog model that has them. Narrow one of those with that model's own `input`. A catalog provider has no `models` list to put it in, so write it under `modelOverrides`, keyed by model id:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
anthropic:
|
||||
modelOverrides:
|
||||
claude-sonnet-4-5:
|
||||
input: [text]
|
||||
```
|
||||
|
||||
Every list must name at least one modality except a model's own, where an empty list means the same as omitting it. An unknown modality is refused wherever it is written.
|
||||
|
||||
Both fields state a claim about your endpoint rather than checking it. A model that declares images its endpoint does not serve is not caught here; the provider rejects the request instead.
|
||||
|
||||
## Select a model
|
||||
|
||||
Configured providers appear in the model picker. Selecting a model also makes it the default for new sessions. A session that has already sent a request retains the model recorded in its own log.
|
||||
@@ -39,6 +90,8 @@ If a saved default names a provider that was deleted, the composer displays **Se
|
||||
- **`MISSING_CREDENTIAL`** — Store the provider key through the Models page or supply the referenced environment variable.
|
||||
- **`UNKNOWN_MODEL`** — Select a configured model or add the missing model to the custom provider.
|
||||
- **Fetching available models returns 401** — Check the key. Model discovery calls the OpenAI-compatible `GET /models` endpoint; enter models manually for endpoints that do not provide it.
|
||||
- **An image is refused before sending** — The model declares no image modality. Give a custom provider's model `input: [text, image]`; DeepSeek's own chat-completions route is text-only and cannot be configured otherwise.
|
||||
- **The provider rejects a request carrying an image** — The model declares images its endpoint does not actually serve. Remove `image` from whichever list granted it — the model's `input`, or the route's `defaultInput` — then start a new session: the attached image stays in the session log, so the same request repeats until the session moves off it.
|
||||
|
||||
## Advanced configuration
|
||||
|
||||
|
||||
@@ -28,6 +28,57 @@ Provider ID 是永久的,因为请求、已保存会话、模型默认值和
|
||||
|
||||
在**模型目录**中选择**获取可用模型**,可查询表单当前显示的基础 URL 和凭据。选择候选项只会更新草稿;保存前不会存储提供方。目录提供方使用已安装目录,不发起网络请求。
|
||||
|
||||
### 图片输入
|
||||
|
||||
手动输入的模型在自己声明之前一律按纯文本对待,因为没有任何环节能去询问端点接受哪些模态。给这类模型附加图片,会在发送前就被拒绝,并点名该模型。
|
||||
|
||||
因此自定义提供方下的视觉模型需要加一行。表单没有对应字段;请在 `$DSH_HOME/settings.yaml` 中给该模型加上 `input`:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
my-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://gateway.example/v1
|
||||
models:
|
||||
- id: legacy-chat
|
||||
- id: vision-preview
|
||||
input: [text, image]
|
||||
```
|
||||
|
||||
`input` 接受 `text` 和 `image`,且只作用于该模型,因此一条路由可以同时服务两类模型。省略它——或写成空列表,两者同义——则保留已安装目录为该模型记录的模态;目录未描述的模型则回退到该路由的 `defaultInput`。
|
||||
|
||||
如果你手动录入的模型全都接受图片,可以在路由上设置一次回退值,不必逐个模型写:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
vision-gateway:
|
||||
apiKeyEnv: GATEWAY_API_KEY
|
||||
api: openai-completions
|
||||
baseURL: https://vision.example/v1
|
||||
defaultInput: [text, image]
|
||||
models:
|
||||
- id: first-model
|
||||
- id: second-model
|
||||
```
|
||||
|
||||
`defaultInput` 是回退值而不是覆盖值,默认为 `[text]`:在目录提供方上,它只为目录未描述的模型作答,因此绝不会把目录中本就具备图片能力的模型的该能力去掉。要收窄这类模型,请用它自己的 `input`。目录提供方没有可供填写的 `models` 列表,因此写在 `modelOverrides` 下,以模型 id 为键:
|
||||
|
||||
```yaml
|
||||
llm-pi-ai:
|
||||
providers:
|
||||
anthropic:
|
||||
modelOverrides:
|
||||
claude-sonnet-4-5:
|
||||
input: [text]
|
||||
```
|
||||
|
||||
除模型自身的列表外,每个列表都至少要写一项模态;模型自身的空列表与省略它同义。未知模态在任何位置写入都会被拒绝。
|
||||
|
||||
这两个字段都是对你端点的断言,而不是对它的检查。声明了端点并不提供的图片能力的模型不会在这里被拦下,改由提供方拒绝该请求。
|
||||
|
||||
## 选择模型
|
||||
|
||||
已配置的提供方会出现在模型选择器中。选择模型也会将其设为新会话的默认值。已发送过请求的会话会保留自身日志中记录的模型。
|
||||
@@ -39,6 +90,8 @@ Provider ID 是永久的,因为请求、已保存会话、模型默认值和
|
||||
- **`MISSING_CREDENTIAL`**:通过模型页存储提供方密钥,或提供被引用的环境变量。
|
||||
- **`UNKNOWN_MODEL`**:选择已配置的模型,或向自定义提供方添加缺失的模型。
|
||||
- **获取可用模型返回 401**:检查密钥。模型发现会调用 OpenAI 兼容的 `GET /models` 端点;对于不提供该端点的服务,请手动输入模型。
|
||||
- **图片在发送前被拒绝**:该模型未声明图片模态。请给自定义提供方的模型加上 `input: [text, image]`;DeepSeek 自身的 chat-completions 路由是纯文本的,且无法通过配置改变。
|
||||
- **提供方拒绝了带图片的请求**:该模型声明了其端点实际并不提供的图片能力。请从授予它图片能力的那个列表中移除 `image`——可能是模型的 `input`,也可能是路由的 `defaultInput`——然后开启新会话:附加的图片会留在会话日志里,因此在会话离开它之前,同一个请求会不断重复。
|
||||
|
||||
## 进阶配置
|
||||
|
||||
|
||||
Reference in New Issue
Block a user