机器翻译正文由机器翻译自项目原始文档(英文),排版经程序统一处理,可能存在偏差,请以原项目仓库为准。
时尚CLIP
快速开始
| 名称 | 链接 |
|---|---|
| FashionCLIP 特征提取与分类 | |
| 教程 - 使用 RecList 评估 FashionCLIP |
更新(23/10/03):我们已更新模型!我们发现 laion/CLIP-ViT-B-32-laion2B-s34B-b79K 检查点(感谢 Bin!)在 Fashion 上的效果优于原始 OpenAI CLIP。因此,我们微调了更新(更好!)的 FashionCLIP 版本(以下简称 FashionCLIP 2.0),同时保持架构不变。我们推测 laion/CLIP-ViT-B-32-laion2B-s34B-b79K 带来的性能提升是由于训练数据增加(是 OpenAI CLIP 数据的 5 倍)。然而,我们的 论文 观点保持不变 —— 在我们的时尚数据集上微调 laion/CLIP 提高了零样本在各基准上的表现。请参见下表,比较各模型的加权宏 F1 分数。
| 模型 | FMNIST | KAGL | DEEP |
|---|---|---|---|
| OpenAI CLIP | 0.66 | 0.63 | 0.45 |
| FashionCLIP | 0.74 | 0.67 | 0.48 |
| Laion CLIP | 0.78 | 0.71 | 0.58 |
| FashionCLIP 2.0 | 0.83 | 0.73 | 0.62 |
我们现在已入驻 Hugging Face!该模型可在 这里 获取。
我们现在已入驻 自然科学报告!
引用
@Article{Chia2022,
title="Contrastive language and vision learning of general fashion concepts",
author="Chia, Patrick John
and Attanasio, Giuseppe
and Bianchi, Federico
and Terragni, Silvia
and Magalh{\~a}es, Ana Rita
and Goncalves, Diogo
and Greco, Ciro
and Tagliabue, Jacopo",
journal="Scientific Reports",
year="2022",
month="Nov",
day="08",
volume="12",
number="1",
pages="18958",
abstract="The steady rise of online shopping goes hand in hand with the development of increasingly complex ML and NLP models. While most use cases are cast as specialized supervised learning problems, we argue that practitioners would greatly benefit from general and transferable representations of products. In this work, we build on recent developments in contrastive learning to train FashionCLIP, a CLIP-like model adapted for the fashion industry. We demonstrate the effectiveness of the representations learned by FashionCLIP with extensive tests across a variety of tasks, datasets and generalization probes. We argue that adaptations of large pre-trained models such as CLIP offer new perspectives in terms of scalability and sustainability for certain types of players in the industry. Finally, we detail the costs and environmental impact of training, and release the model weights and code as open source contribution to the community.",
issn="2045-2322",
doi="10.1038/s41598-022-23052-9",
url="https://doi.org/10.1038/s41598-022-23052-9"
}
信息
我们正在等待 Farfetch 数据集的官方发布,届时经过微调的模型权重、预处理的图像和文本向量将会公开。同时,我们目前使用 Hugging Face 对 CLIP 的实现,并可以通过遵循标准的 Hugging Face 命名规范(即 fclip = FashionCLIP('<username>/<repo_name>', ... ) )使用来自 OpenAI 的模型权重。我们也支持私有仓库(即 fclip = FashionCLIP('<username>/<repo_name>', auth_token=<AUTH_TOKEN>, ... ) )。
详细信息见下文!
概览
FashionCLIP 是一个类似 CLIP 的模型,经过针对时尚行业的微调。我们对 CLIP (Radford 等, 2021) 在来自 Farfetch 数据集的 70 万对以上样本上进行了微调[1]。
我们通过将 FashionCLIP 应用于行业中的实际问题,如检索、分类和时尚解析,来评估其性能。我们的结果表明,微调有助于捕捉特定领域的概念,并在零样本场景中进行泛化。我们还通过定性分析来补充定量测试,并提供了初步见解,说明如何在视觉空间中建立的概念能够解锁语言泛化。更多详情请参阅我们的论文。
在此仓库中,你将找到用于与 FashionCLIP 交互的 API 以及使用 streamlit 构建的交互演示(即将推出!),展示了 FashionCLIP 的功能。
API 与演示
快速使用指南
需要快速生成嵌入向量吗?想测试检索性能吗?
首先,你应该能够通过 pip 快速安装此工具。
$ pip install fashion-clip
如果你有文本列表和图像路径,生成嵌入非常容易:
from fashion_clip.fashion_clip import FashionCLIP
fclip = FashionCLIP('fashion-clip')
# we create image embeddings and text embeddings
image_embeddings = fclip.encode_images(images, batch_size=32)
text_embeddings = fclip.encode_text(texts, batch_size=32)
# we normalize the embeddings to unit norm (so that we can use dot product instead of cosine similarity to do comparisons)
image_embeddings = image_embeddings/np.linalg.norm(image_embeddings, ord=2, axis=-1, keepdims=True)
text_embeddings = text_embeddings/np.linalg.norm(text_embeddings, ord=2, axis=-1, keepdims=True)
使用我们的 colab 笔记本查看更多功能。
HF API
from PIL import Image
import requests
from transformers import CLIPProcessor, CLIPModel
model = CLIPModel.from_pretrained("patrickjohncyh/fashion-clip")
processor = CLIPProcessor.from_pretrained("patrickjohncyh/fashion-clip")
image = Image.open("images/image1.jpg")
inputs = processor(text=["a photo of a red shoe", "a photo of a black shoe"],
images=image, return_tensors="pt", padding=True)
outputs = model(**inputs)
logits_per_image = outputs.logits_per_image # this is the image-text similarity score
probs = logits_per_image.softmax(dim=1)
print(probs)
image.resize((224, 224))
额外内部 FashionCLIP API
安装
从项目根目录,本地安装 fashion-clip 包,使用
$ pip install -e .
有两个主要的抽象用于方便使用 FashionCLIP。
首先是 FCLIPDataset 类,它封装了与给定目录相关的信息,并公开了对 FashionCLIP 至关重要的信息。此外,它还提供了用于快速探索和可视化数据的辅助函数。主要的初始化参数是
name: str -> Name of dataset
image_source_path: str -> absolute path to images (can be local or s3)
image_source_type: str -> type of source (i.e. local or s3)
catalog: List[dict] = None -> list of dicts containing at miniumum the keys ['id', 'image', 'caption']
为了方便使用,API 还提供对数据集(一旦正式发布)的访问,该数据集在论文中用于训练 FahionCLIP,只需指定相应的目录名称即可。
预先包含的数据集
from fashion_clip import FCLIPDataset
dataset = FCLIPDataset(name='FF',
image_source_path='path/to/images',
image_source_type='local')
自定义数据集
from fashion_clip import FCLIPDataset
my_catalog = [{'id': 1, 'image': 'x.jpg', 'caption': 'image x'}]
dataset = FCLIPDataset(name='my_dataset',
image_source_path='path/to/images',
image_source_type='local',
catalog=my_catalog)
第二个抽象是 FashionCLIP 类,它接受一个 Hugging Face CLIP 模型名称和一个 FCLIPDataset,并提供方便的功能来执行多模态检索、零样本分类和定位等任务。FashionCLIP 的初始化参数如下:
model_name: str -> Name of model OR path to local model
dataset: FCLIPDataset -> Dataset,
normalize: bool -> option to convert embeddings to unit norm
approx: bool -> option to use approximate nearest neighbors
类似于 FCLIPDataset 抽象,我们已经包含了来自论文的预训练 FashionCLIP 模型,托管在 这里。如果收到未知的数据集和模型组合,将在对象实例化时生成图像和标题向量,否则将从 S3 中提取预先计算的向量/嵌入。
from fashion_clip import FCLIPDataset, FashionCLIP
dataset = FCLIPDataset(name='FF',
image_source_path='path/to/images',
image_source_type='local')
fclip = FashionCLIP('fasihon-clip', ff_dataset)
有关如何使用该软件包的更多详细信息,请参阅随附的笔记本!
有趣的相关项目!
- 查看 RustEmbed 以获取使用 gRPC 创建 FashionCLIP 嵌入的应用程序。
注释:
[1] 待官方发布。↑
- 本文标题:fashion-clip - FashionCLIP 是一个类似 CLIP 的模型
- 本文链接:https://www.cn121.com/shop/patrickjohncyh-fashion-clip.html
- 原项目:patrickjohncyh/fashion-clip 版权归原作者 patrickjohncyh 及贡献者所有
- 收录信息:本站于 2026-10-09 收录本项目,本页所列协议与仓库指标均为收录当时的状态;该日期之后原项目的版本更新与协议变更,本页不作同步。
- 开源协议:收录时本项目采用 MIT(查看 LICENSE 原文),本站译文为其衍生内容;使用、修改、分发请以该仓库 LICENSE 原文为准。
- 站点出处:本文首发于 OneTwoOne,收录自 GitHub 开源项目 patrickjohncyh/fashion-clip。
- 翻译说明:本页正文为人工智能生成内容——由机器翻译对原项目 README 初译、经程序校验排版,可能存在错漏,请以原项目文档为准。
- 引用声明:商业转载、第三方聚合或 AI 检索训练引用时,请务必保留以上来源出处、本文永久链接,以及原项目的版权声明与许可信息。
- 下架通道:若原项目此后变更或收紧了许可协议、或作者/权利人认为本站的收录方式(译文、排版适配、简介翻译等)超出其授权范围,请通过 xyd3302001@163.com 发送下架通知,并附上项目地址与本页链接。本站核实后将第一时间删除本页内容,或改为不复制原文的目录性收录;署名更正等其他要求可一并提出。