---
title: "Translating Visuals into Words: Image Captioning with AI"
url: "https://maker.wiznet.io/Benjamin/projects/translating-visuals-into-words%3A-image-captioning-with-ai/"
markdown_url: "https://maker.wiznet.io/Benjamin/projects/translating-visuals-into-words%3A-image-captioning-with-ai/md"
type: "WCC: WIZnet Created Content"
author: "Benjamin"
author_url: "https://maker.wiznet.io/Benjamin/"
original_author: "benjamin"
published: "2023-08-10"
language: "en"
hardware: ["WIZnet W5100S-EVB-Pico", "Arducam HM0360 Camera Module"]
likes: 4
views: 3651
comments: 1
source: "WIZnet Makers (https://maker.wiznet.io/)"
---

# Translating Visuals into Words: Image Captioning with AI

> basic project conducted with the W5100S-EVB-PICO board, utilizing a replicated image captioning model to generate descriptions for images

Original author: benjamin

## Components

- **WIZnet W5100S-EVB-Pico** x 2 ([docs](https://docs.wiznet.io/Product/Chip/Ethernet/W5100S/w5100s-evb-pico))
- **Arducam HM0360 Camera Module** x 1
- Software: **Adafruit Circuitpython** ([docs](https://docs.circuitpython.org/en/latest/README.html))
- Software: **MicroPython** ([docs](http://micropython.org/))

## Documents and links

- [circuitpython - arducam](https://github.com/wiznetmaker/W5100S-EVB-PICO-ImageCaptioning/tree/main/pico_circuitpython_arducam) (code): I modified the code a bit to get a more stable connection.
- [micropython-captioning](https://github.com/wiznetmaker/W5100S-EVB-PICO-ImageCaptioning/tree/main/pico_micropython_ssd1306) (code): Code using the API for IMG2TXT captioning

## Article

## Overview

In this project, we are using two W5100S-EVB-PICO boards.

1.The first board connects an Arducam and Ethernet to serve the role of transmitting a picture to a web page upon receiving a web request.

[![image](https://user-images.githubusercontent.com/115054808/260601615-d8921838-cfed-4157-ab0b-ed93ded69172.png)](https://user-images.githubusercontent.com/115054808/260601615-d8921838-cfed-4157-ab0b-ed93ded69172.png)

2.The second board will perform image-to-text captioning via the "Replicate API" in the form of a web address serving images from the first PICO board over an Ethernet connection and display them on the ssd1306 OLED screen.

[![image](https://user-images.githubusercontent.com/115054808/260602763-71c9cca9-484c-4927-896a-577ff6c80bb8.png)](https://user-images.githubusercontent.com/115054808/260602763-71c9cca9-484c-4927-896a-577ff6c80bb8.png)

Discuss this in more detail below.

## Model used for image captioning

BLIP-2 is a part of Salesforce's LAVIS project. BLIP-2 is a generic and efficient pre-training strategy that leverages the advancements of pretrained vision models and large language models (LLMs). BLIP-2 surpasses Flamingo in zero-shot VQAv2 (scoring 65.0 vs 56.3) and sets a new state-of-the-art in zero-shot captioning (achieving a CIDEr score of 121.6 on NoCaps compared to the previous best of 113.2). When paired with powerful LLMs such as OPT and FlanT5, BLIP-2 unveils new zero-shot instructed vision-to-language generation capabilities for a range of intriguing applications.

[![image](https://user-images.githubusercontent.com/115054808/260513344-424c38c5-62f4-48fb-8cbc-565dfafbcffe.png)](https://user-images.githubusercontent.com/115054808/260513344-424c38c5-62f4-48fb-8cbc-565dfafbcffe.png)

[Model testable on Replicate](https://replicate.com/andreasjansson/blip-2)

[![image](https://user-images.githubusercontent.com/115054808/260513161-8bd98582-f653-4169-8124-f485d8cdfb1d.png)](https://user-images.githubusercontent.com/115054808/260513161-8bd98582-f653-4169-8124-f485d8cdfb1d.png)

[Research paper on the model](https://arxiv.org/abs/2301.12597)

[Official GitHub](https://github.com/salesforce/LAVIS/tree/main/projects/blip2)

## Installation and How to Use

in progress

## Outputs

[![image](https://user-images.githubusercontent.com/115054808/260603573-ba67ff4a-165a-42aa-ae7d-bbc36020d29b.png)](https://user-images.githubusercontent.com/115054808/260603573-ba67ff4a-165a-42aa-ae7d-bbc36020d29b.png)

[![image](https://user-images.githubusercontent.com/115054808/260603914-9ae158fd-2eee-489d-ac73-2792e02c3d0d.png)](https://user-images.githubusercontent.com/115054808/260603914-9ae158fd-2eee-489d-ac73-2792e02c3d0d.png)

## Potential Developments

The current focus of this project is on captioning an image and displaying the text. However, by adding speakers in the future, this could evolve into applications for the visually impaired or anomaly detection, among various other project ideas.

## Contributions and Feedback

Contributions and feedback on the project are always welcome. Please use the GitHub issue tracker to report issues or create pull requests.

## Original link

[WIZnet maker Github](https://github.com/wiznetmaker/W5100S-EVB-PICO-ImageCaptioning)

---

Source: https://maker.wiznet.io/Benjamin/projects/translating-visuals-into-words%3A-image-captioning-with-ai/
