ttengwang/Caption-Anything
A multi-model image captioning tool that combines Segment Anything, visual captioning, and ChatGPT for controllable text generation.

Not currently ranked — collecting fresh signals.
star history
Caption-Anything is an image processing system that leverages Segment Anything for visual segmentation, a captioning model for text generation, and ChatGPT for language-level control. Users can click on image regions to select objects, then generate descriptive captions with customizable style, length, sentiment, and factuality. The system also supports conversational follow-up about selected objects via ChatGPT integration.
Frequently asked
- What is ttengwang/Caption-Anything?
- A multi-model image captioning tool that combines Segment Anything, visual captioning, and ChatGPT for controllable text generation.
- Is Caption-Anything open source?
- Yes — ttengwang/Caption-Anything is open source, released under the BSD-3-Clause license.
- What language is Caption-Anything written in?
- ttengwang/Caption-Anything is primarily written in Python.
- How popular is Caption-Anything?
- ttengwang/Caption-Anything has 1.8k stars on GitHub.
- Where can I find Caption-Anything?
- ttengwang/Caption-Anything is on GitHub at https://github.com/ttengwang/Caption-Anything.