Proxy Lite 模型卡

这是 Proxy 的小型开源权重版本
模型描述
- 开发方: Convergence AI
- 模型类型: 3B 视觉-语言模型(Vision-Language Model)
- 智能体类型: 网页浏览智能体(Web-browsing Agent)
- 许可证: CC-BY-NC-4.0
- 微调自: Qwen/Qwen2.5-VL-3B-Instruct
- 运行该智能体
在浏览器中运行 Proxy Lite
请访问 https://github.com/convergence-ai/proxy-lite 以在浏览器中运行 Proxy Lite。
git clone https://github.com/convergence-ai/proxy-lite.git
make proxy
proxy "Find some markets near Kings Cross and tell me their ratings."

用途
Proxy Lite 专为在网页浏览器中完成自动化任务而设计和训练。
运行该模型的完整代码已在 GitHub 仓库 中开源,其中包含用于运行模型的 CLI 工具以及一个 Streamlit 应用。
你可以使用此 端点 进行小规模测试。
直接使用
我们推荐使用 vLLM 自托管你的端点,你可以使用以下命令:
vllm serve convergence-ai/proxy-lite-3b \
--trust-remote-code \
--enable-auto-tool-choice \
--tool-call-parser hermes \
--port 8008 \
上述工具参数对于正确解析模型生成的工具调用非常关键。
重要提示: 最新版的
transformers尚未支持 Qwen-2.5-VL,因此请务必从源码安装。
消息历史
关于使用与提示 Proxy Lite 的更多细节,请参考其 代码仓库,但该模型期望的消息历史格式大致如下:
message_history = [
{
"role": "system",
"content": "You are Proxy Lite...", # 完整系统提示见 src/proxy_lite/agents/proxy_lite_agent.py
}, # 系统提示
{
"role": "user",
"content": "Find some markets near Kings Cross and tell me their ratings.",
}, # 设定任务
{
"role": "user",
"content": [
{"type": "image_url", "image_url": {base64_encoded_screenshot} },
{"type": "text", "text": "URL: https://www.google.com/ \n- [0] <a>About</a> \n- [1] <a>Store</a>...."}
] # 来自环境的观察信息
},
]
如此便能不断累积消息历史,交替在助手(assistant,提供动作)与用户(user,提供观察)之间。
上下文窗口管理: 在调用模型时,除当前观察之外的所有历史观察都会被丢弃,以减少图像 token 的巨大开销。由于模型回复中已包含对先前观察的反思,并会全部纳入消息历史,因此模型在规划新动作时仍能感知到完整的历史信息。
工具
你还应当传入模型可访问的 Tools,它们定义了模型可用的动作空间。使用 transformers 时可参考以下方式:
from qwen_vl_utils import process_vision_info
from transformers import AutoProcessor
from proxy_lite.tools import ReturnValueTool, BrowserTool
from proxy_lite.serializer import OpenAICompatableSerializer
processor = AutoProcessor.from_pretrained("convergence-ai/proxy-lite-3b")
tools = OpenAICompatableSerializer().serialize_tools([ReturnValueTool(), BrowserTool(session=None)])
templated_messages = processor.apply_chat_template(
message_history, tokenize=False, add_generation_prompt=True, tools=tools
)
image_inputs, video_inputs = process_vision_info(message_history)
batch = processor(
text=[templated_messages],
images=image_inputs,
videos=video_inputs,
padding=True,
return_tensors="pt",
)
或者直接向端点发起请求,由服务端处理格式:
from openai import OpenAI
client = OpenAI(base_url="http://convergence-ai-demo-api.hf.space/v1")
response = client.chat.completions.create(
model="convergence-ai/proxy-lite-3b",
messages=message_history,
tools=tools,
tool_choice="auto",
)
评测
Proxy Lite 在 WebVoyager 基准测试中取得了 72.4% 的成绩,在所有可用的开源权重模型中排名第一。
按网站细分的结果如下:
| 网站名称(web_name) | 成功率 (%) | 完成率 (%) | 平均步骤数 |
|---|---|---|---|
| Allrecipes | 87.8 | 95.1 | 10.3 |
| Amazon | 70.0 | 90.0 | 7.1 |
| Apple | 82.1 | 89.7 | 10.7 |
| ArXiv | 60.5 | 79.1 | 16.0 |
| BBC News | 69.4 | 77.8 | 15.9 |
| Booking | 70.0 | 85.0 | 24.8 |
| Cambridge Dict. | 86.0 | 97.7 | 5.7 |
| Coursera | 82.5 | 97.5 | 4.7 |
| ESPN | 53.8 | 87.2 | 14.9 |
| GitHub | 85.0 | 92.5 | 10.0 |
| Google Flights | 38.5 | 51.3 | 34.8 |
| Google Map | 78.9 | 94.7 | 9.6 |
| Google Search | 71.4 | 92.9 | 6.0 |
| Huggingface | 68.6 | 74.3 | 18.4 |
| Wolfram Alpha | 78.3 | 93.5 | 6.1 |
超出范围的使用
Proxy Lite 专为在浏览器环境中自动化执行日常任务而设计。然而,它不应被用于以下场景:
- 高风险或安全关键型应用:
请勿将 Proxy Lite 用于金融交易、医疗操作、法律决策或应急响应等任务,因为这些场景中任何错误都可能导致严重伤害或重大经济损失。 - 未授权或具有侵入性的数据采集:
对网站进行自动化抓取或数据提取只能在获得明确许可的情况下进行。Proxy Lite 不应被用于绕过网站的服务条款、版权限制或隐私政策。 - 与恶意或未经验证网站的交互:
使用该模型浏览或交互可疑、不可信的网站,可能会使系统面临恶意软件、钓鱼攻击或其他网络攻击等安全威胁。 - 受合规监管或法律敏感的操作:
需要严格遵守法律或监管标准的任务(例如处理个人数据或敏感信息)应在本模型之外额外采取保护措施。
引用
BibTeX:
@article{proxy-lite,
title={Proxy Lite - A Mini, Open-weights, Autonomous Assistant},
author={Convergence AI},
year={2025}
}
模型树:convergence-ai/proxy-lite-3b
- 基础模型:Qwen/Qwen2.5-VL-3B-Instruct — 微调而来(另有 877 个 派生模型)
- 本模型
- 微调派生:1 个模型
- 量化版本:2 个模型