Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,14 @@

## [Unreleased]

- **看图把关的候选编号打架**:提示词里写"候选 0:[图片1]"——候选从 0 起、引擎留的图片占位从 1 起,模型回的 `i` 按哪个数没法确定。
本机 3B 视觉模型真机 3 选 1 只对 1/5:第 1 条候选永远被算成 0 分。现在候选也从 1 起编号,解析兼容仍按 0 起数的模型(出现 i=0 就按 0 起)。
同一模型同一批图:1/5 → 4/5(剩下那条是把猫雕像当真猫,模型的极限)。
- **看图把关可以走本机 Ollama 的视觉模型**(qwen2.5vl / llava / minicpm-v…;型号只列名字带 vl / llava / vision 的),至此写脚本 + 看图 + 出图
三步都能 0 key 全本机。**需要引擎 ≥ 0.19.3**(AO 的 ollama 连接器以前剥掉图片,已改为走它的 `images` 字段并把图片算进 num_ctx——
真机 12 张缩略图只按文本估 4096,Ollama 直接 400);旧引擎下"验证并开启"真发一张红图,会如实报看不了图。
真机 3B:一镜从"本机出图 56 秒"变成"本机看图选中素材 44 秒",但它给了一张满是文字的示意图 7 分(提示词要求 ≤ 2)——3B 判断力有限,界面提示建议 7B。

- **MCP server(`openshorts mcp`)**:让 Claude Code / Cursor 等 agent 直接"话题 → 成片"。五个工具:`create_video`、`render_project`、`job_status`、
`list_projects`、`doctor`。出片 1–25 分钟远超工具调用超时,所以任务化:立刻回任务号,状态落盘在 `~/.openshorts/mcp-jobs/`;
server 被客户端重启后旧任务如实报 interrupted,不让 agent 永远等;子进程不 detach,server 退出一并收掉。任务跑的是 CLI 的
Expand Down
4 changes: 3 additions & 1 deletion server/kaipian.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -278,7 +278,9 @@ kaipian.get('/providers/text', async (_req, res, next) => {
}));
// 本机 Ollama 不在 AO 的 API 供应商表里(它不要 key),单独探测后排在最前:不花钱的放前面
const ol = await ollamaStatus();
list.unshift({ id: 'ollama', local: true, running: ol.running, baseUrl: ol.baseUrl, hasKey: ol.running && ol.models.length > 0, fromEnv: false, envKey: null, models: ol.models, visionModels: [], vision: false });
// 看图把关也能走本机:只列名字带 vl / llava / vision 的型号(需要引擎 ≥ 0.19.3 才真的把图发给 Ollama;旧引擎会剥图,
// "验证并开启"那一步真发一张红图,答不出来就会如实报"这个模型看不了图")
list.unshift({ id: 'ollama', local: true, running: ol.running, baseUrl: ol.baseUrl, hasKey: ol.running && ol.models.length > 0, fromEnv: false, envKey: null, models: ol.models, visionModels: ol.visionModels ?? [], vision: ol.running && (ol.visionModels ?? []).length > 0 });
const c = readConfig();
res.json({ providers: list, vision: c.vision ?? { provider: '', model: '' }, text: c.text ?? { provider: '', model: '' } });
} catch (e) { next(e); }
Expand Down
3 changes: 2 additions & 1 deletion src/kaipian/Kaipian.tsx
Original file line number Diff line number Diff line change
Expand Up @@ -342,8 +342,9 @@ export const Kaipian = () => {
<label>{t('供应商')}
<select value={vis.provider} onChange={(e) => setVis({provider: e.target.value, model: (textProv?.providers.find((p) => p.id === e.target.value)?.visionModels?.[0]) ?? ''})}>
<option value="">{t('不开')}</option>
{(textProv?.providers ?? []).filter((p) => p.vision).map((p) => <option key={p.id} value={p.id}>{p.id}{p.hasKey ? ' ✓' : ''}</option>)}
{(textProv?.providers ?? []).filter((p) => p.vision).map((p) => <option key={p.id} value={p.id}>{p.local ? t('ollama(本机,免费,不用 key)') : p.id}{p.hasKey ? ' ✓' : ''}</option>)}
</select></label>
{vis.provider === 'ollama' && <p className="kp-hint">{t('用本机视觉模型看图:不花钱、不联网。3B 的小模型 3 选 1 能对 4/5,偶尔把雕像当真猫;想更稳装 7B:')}<code>ollama pull qwen2.5vl:7b</code></p>}
{vis.provider && <label>{t('模型')}<input list="kp-vm" value={vis.model} onChange={(e) => setVis({...vis, model: e.target.value})} placeholder={t('模型 id(可手填)')}/></label>}
<datalist id="kp-vm">{(textProv?.providers.find((p) => p.id === vis.provider)?.visionModels ?? []).map((m) => <option key={m} value={m}/>)}</datalist>
<button className="primary" onClick={saveVision} disabled={!!busy}>{vis.provider ? t('验证并开启') : t('关闭看图把关')}</button>
Expand Down
1 change: 1 addition & 0 deletions src/kaipian/i18n.ts
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ const EN: Record<string, string> = {
'没连上本机的 Ollama(': 'Cannot reach Ollama on this machine (', ')。装好并启动后回来刷新:': '). Install and start it, then refresh: ',
'Ollama 在跑,但还没装能写稿的模型。终端里跑:': 'Ollama is running but has no chat model yet. In a terminal run: ',
'用这台机器上的模型写脚本:不花钱、不联网。7B 级的小模型偶尔写得偏短,开片会自动要求重写一次;想更稳就换 14B 以上。': 'Write scripts with a model on this machine: free and offline. 7B-class models sometimes write too short — OpenShorts asks for one rewrite automatically; use 14B+ for steadier results.',
'用本机视觉模型看图:不花钱、不联网。3B 的小模型 3 选 1 能对 4/5,偶尔把雕像当真猫;想更稳装 7B:': 'Judge footage with a vision model on this machine: free and offline. A 3B model picks the right clip 4 times out of 5 and occasionally mistakes a statue for a cat; for steadier results install 7B: ',
'复制诊断信息': 'Copy diagnostics', '已复制 ✓': 'Copied ✓', '正在体检…': 'Checking…', '去反馈 ↗': 'Report an issue ↗',
'报问题时贴上它:版本、系统、体检结果,key 已打码。': 'Paste this when reporting a problem: version, system, health check — keys are masked.',
'还没有写脚本用的文本模型 key。': 'No text-model key for script writing yet. Open ',
Expand Down
9 changes: 6 additions & 3 deletions src/local/ollama.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -13,14 +13,17 @@ export const ollamaBaseUrl = (env = process.env) => {
/** 嵌入模型聊不了天:列进下拉等于埋雷(选了它写脚本直接报错) */
const isEmbedding = (m) => /embed|bge-|minilm|nomic|e5-|gte-/i.test(m.name ?? '') || /bert/i.test(m.details?.family ?? '');

/** 能看图的本机模型(按名字判:qwen2.5vl / llava / minicpm-v / moondream / llama3.2-vision…)——看图把关只能选这些 */
export const isVisionModel = (name) => /vl|llava|vision|minicpm-v|moondream|bakllava|gemma3/i.test(String(name));

export async function ollamaStatus({ fetchImpl = fetch, env = process.env, timeoutMs = 1500 } = {}) {
const baseUrl = ollamaBaseUrl(env);
try {
const r = await fetchImpl(`${baseUrl}/api/tags`, { signal: AbortSignal.timeout(timeoutMs) });
if (!r.ok) return { running: false, baseUrl, models: [], reason: `HTTP ${r.status}` };
if (!r.ok) return { running: false, baseUrl, models: [], visionModels: [], reason: `HTTP ${r.status}` };
const all = (await r.json()).models ?? [];
// 大的在前:同一台机器上,参数多的那个写口播稿明显更稳
const models = all.filter((m) => !isEmbedding(m)).sort((a, b) => (b.size ?? 0) - (a.size ?? 0)).map((m) => m.name);
return { running: true, baseUrl, models };
} catch (e) { return { running: false, baseUrl, models: [], reason: String(e?.cause?.code ?? e?.name ?? e).slice(0, 60) }; }
return { running: true, baseUrl, models, visionModels: models.filter(isVisionModel) };
} catch (e) { return { running: false, baseUrl, models: [], visionModels: [], reason: String(e?.cause?.code ?? e?.name ?? e).slice(0, 60) }; }
}
15 changes: 12 additions & 3 deletions src/sources/rank.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,16 @@ export async function evidenceFrame(candidate, { fetchImpl = fetch, timeoutMs =
/** 解析回复里的 JSON 数组 [{i, score, why}] */
export function parseScores(text, n) {
const m = String(text).match(/\[[\s\S]*\]/); if (!m) return null;
try { const arr = JSON.parse(m[0]); const out = Array.from({ length: n }, (_, i) => ({ i, score: 0, why: '' })); for (const x of arr) { const i = Number(x.i ?? x.index); if (i >= 0 && i < n) out[i] = { i, score: Math.max(0, Math.min(10, Number(x.score) || 0)), why: String(x.why ?? '') }; } return out; } catch { return null; }
try {
const arr = JSON.parse(m[0]); if (!Array.isArray(arr)) return null;
// 候选按 1..n 编号(与连接器留的 [图片N] 占位一致——以前"候选 0:[图片1]"两套编号打架,模型回的 i 按哪个数没法确定)。
// 兼容仍按 0 起数的模型:出现了 i=0 或 i=n 越界的情况就按 0 起解释
const idx = arr.map((x) => Number(x.i ?? x.index)).filter(Number.isFinite);
const oneBased = !idx.includes(0) && idx.every((i) => i >= 1 && i <= n);
const out = Array.from({ length: n }, (_, i) => ({ i, score: 0, why: '' }));
for (const x of arr) { const i = Number(x.i ?? x.index) - (oneBased ? 1 : 0); if (i >= 0 && i < n) out[i] = { i, score: Math.max(0, Math.min(10, Number(x.score) || 0)), why: String(x.why ?? '') }; }
return out;
} catch { return null; }
}

/**
Expand All @@ -89,8 +98,8 @@ export async function rankCandidates(candidates, intent, { connector, cfg, thres
if (!usable.length) return candidates.map((c) => ({ ...c, score: null }));
const zh = /[一-鿿]/.test(intent);
const prompt = (zh
? [`你是短视频剪辑师。下面是同一段口播要配的画面意图,以及 ${usable.length} 条候选素材各一帧。给每条打分 0–10:画面主体、场景与意图是否匹配(主体对得上给 6 分起,完全无关 0–2 分,图表/文字/标题卡一律 ≤ 2)。`, `画面意图:${intent}`, ...usable.map((x, k) => `候选 ${k}:${x.f}`), '只输出 JSON 数组:[{"i":0,"score":7,"why":"一句话"}, …]']
: [`You are a video editor. Below is the visual intent for one narration segment and one frame from each of ${usable.length} candidate clips. Score each 0–10 for how well subject/scene match the intent (subject matches → ≥6; unrelated → 0–2; charts/text/title cards ≤ 2).`, `Intent: ${intent}`, ...usable.map((x, k) => `Candidate ${k}: ${x.f}`), 'Output only a JSON array: [{"i":0,"score":7,"why":"…"}, …]']).join('\n');
? [`你是短视频剪辑师。下面是同一段口播要配的画面意图,以及 ${usable.length} 条候选素材各一帧。给每条打分 0–10:画面主体、场景与意图是否匹配(主体对得上给 6 分起,完全无关 0–2 分,图表/文字/标题卡一律 ≤ 2)。`, `画面意图:${intent}`, ...usable.map((x, k) => `候选 ${k + 1}:${x.f}`), `只输出 JSON 数组,i 是候选编号(1 到 ${usable.length}):[{"i":1,"score":7,"why":"一句话"}, …]`]
: [`You are a video editor. Below is the visual intent for one narration segment and one frame from each of ${usable.length} candidate clips. Score each 0–10 for how well subject/scene match the intent (subject matches → ≥6; unrelated → 0–2; charts/text/title cards ≤ 2).`, `Intent: ${intent}`, ...usable.map((x, k) => `Candidate ${k + 1}: ${x.f}`), `Output only a JSON array where i is the candidate number (1 to ${usable.length}): [{"i":1,"score":7,"why":"one sentence"}, …]`]).join('\n');
let scores = null;
// 推理模型(Agnes 2.0-flash)会先吐几百字思考再给 JSON:预算给足。
// 接口偶尔会抽(真机上六镜里抽了一次),所以多试两次并退避——一次失败就等于这一镜没人把关。
Expand Down
5 changes: 4 additions & 1 deletion tests/ollama.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -9,13 +9,16 @@ const TAGS = { models: [
{ name: 'mxbai-embed-large:latest', size: 669_000_000, details: { family: 'bert' } },
{ name: 'qwen2.5:14b', size: 8_988_000_000, details: { family: 'qwen2' } },
{ name: 'qwen2.5-coder:7b', size: 4_683_000_000, details: { family: 'qwen2' } },
{ name: 'qwen2.5vl:3b', size: 3_200_000_000, details: { family: 'qwen25vl' } },
{ name: 'llava:7b', size: 4_700_000_000, details: { family: 'llama' } },
] };
const ok = (body) => async () => ({ ok: true, json: async () => body });

test('嵌入模型不进下拉(选了它写脚本会直接报错);大模型排前面', async () => {
const s = await ollamaStatus({ fetchImpl: ok(TAGS), env: {} });
assert.equal(s.running, true);
assert.deepEqual(s.models, ['qwen2.5:14b', 'qwen2.5-coder:7b', 'llama3:latest']);
assert.deepEqual(s.models, ['qwen2.5:14b', 'llava:7b', 'qwen2.5-coder:7b', 'llama3:latest', 'qwen2.5vl:3b']);
assert.deepEqual(s.visionModels, ['llava:7b', 'qwen2.5vl:3b'], '看图把关只列能看图的;纯文本模型选了等于没开');
});

test('没在跑 / 回了错误码:running=false 并带原因,不抛', async () => {
Expand Down
10 changes: 10 additions & 0 deletions tests/rank.test.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -92,3 +92,13 @@ test('低于补位线的照样判退,绝不混进成片', { skip: !hasFfmpeg &
assert.ok(r.every((c) => !c.fillOnly));
fs.rmSync(d, { recursive: true, force: true });
});

test('parseScores:候选按 1..n 编号时按 1 起解释;模型仍按 0 起(出现 i=0)时按 0 起——真机 3B 视觉模型回 i=1,2 把第 1 条永远算成 0 分', () => {
const one = parseScores('[{"i":1,"score":8},{"i":2,"score":1},{"i":3,"score":2}]', 3);
assert.deepEqual(one.map((x) => x.score), [8, 1, 2]);
const zero = parseScores('[{"i":0,"score":8},{"i":1,"score":1},{"i":2,"score":2}]', 3);
assert.deepEqual(zero.map((x) => x.score), [8, 1, 2]);
// 只回了一部分且从 1 起:缺的那条 0 分,不能把第 1 条的分挪到第 0 条
const partial = parseScores('[{"i":2,"score":9}]', 3);
assert.deepEqual(partial.map((x) => x.score), [0, 9, 0]);
});
Loading