Skip to content

Translating Visuals into Words: Image Captioning with AI

basic project conducted with the W5100S-EVB-PICO board, utilizing a replicated image captioning model to generate descriptions for images

Benjamin

Published August 10, 2023

Translating Visuals into Words: Image Captioning with AI

Components

Hardware components
HM0360 Camera Module

x 1

Software Apps and online services

Project description

Overview In this project, we are using two W5100S-EVB-PICO boards. 1.The first board connects an Arducam and Ethernet to serve the role of transmitting a picture to a web page upon receiving a web request. 2.The second board will perform image-to-text captioning via the "Replicate API" in the form of a web address serving images from the first PICO board over an Ethernet connection and display them on the ssd1306 OLED screen. Discuss this in more detail below. Model used for image captioning BLIP-2 is a part of Salesforce's LAVIS project. BLIP-2 is a generic and efficient pre-training…

Keep reading with a free account

Guests can read 15 projects per visit. Log in or create a free account to read the full description, open the documents and join the comments.