# Visual Search — Evolving search for a new modality

_Senior Design Lead, Microsoft ・ 2018_

Visual Search for Microsoft Bing — evolving search for the camera, from image to knowledge, and growing it into a platform of skills.

---

Given that the world is predominantly visual, search experiences should mirror this visual nature. Visual Search turned the camera into a way to find information when words fall short. I led design and strategy for the effort — building the foundation, envisioning the experiences, and partnering with teams across Microsoft to take it from a hypothesis to a platform of skills used across the search ecosystem.

## Top user problems from UXR

It was **difficult to find similar things on the web**, and **unintuitive to find certain things by searching with text** — the gap was widest exactly where people wanted to know more about the world around them.

![The camera pointed at the Taj Mahal, with a Landmark label pinned to the dome.](https://varunv.com/media/2018/visual-search/image-section-1.webp)

## Goals

A search experience which enables users to find information when words fall short — going beyond the search box. Educate users about a new way to search, and make it top of mind.

![The camera resolving a scene into recognised objects ready to search.](https://varunv.com/media/2018/visual-search/image-section-2.webp)

## Abstracting the experience — how do we go from image to knowledge

### Looking

- User opens the visual search feature
- App begins looking for objects to identify
- Image acquisition, live or static

### Recognizing

- User finds an object to search for
- App identifies objects and begins visual search

### Communicating

Information reveal — based upon recognised information in the image, bounding boxes, hotspots, available modes and flyouts used to reveal top level info and actions.

## Key UX principles

### Digestibility

Show content needed to satisfy user intent and not overwhelming the experience with everything we could possibly show

### Trust & confidence

Let the user know when we know for sure versus when we are less confident to build trust in our product

### Recourse

The user should have option to recourse if the results aren't satisfactory. Minimize failure states — guide the user along a path to successful and relevant results

### Inform the click

The user should know what each click or tap will lead to in a consistent manner

## Visual search invoking flow

The camera had to feel like a natural extension of search — reachable from the feed, and recognizing an object simply by hovering over it.

![The feed, with the camera offered alongside the search box.](https://varunv.com/media/2018/visual-search/section-1.webp)

![The camera opening over the scene in front of the user.](https://varunv.com/media/2018/visual-search/section-2.webp)

![The camera settling on an object and marking it.](https://varunv.com/media/2018/visual-search/section-3.webp)

![The recognised object named, with actions offered.](https://varunv.com/media/2018/visual-search/section-4.webp)

## Camera UI

Focus on an object, and the interface reveals what it is alongside the actions worth taking — look alikes, similar images, and places to shop.

![The camera framing an object, modes along the bottom.](https://varunv.com/media/2018/visual-search/section-5.webp)

![The object identified with a label above it.](https://varunv.com/media/2018/visual-search/section-6.webp)

![Look alikes and similar images offered for the object.](https://varunv.com/media/2018/visual-search/section-7.webp)

![Places to shop the object, listed under the camera view.](https://varunv.com/media/2018/visual-search/section-8.webp)

## Landing the core interaction model

We compared the camera with mode labels, without mode labels, and a universal model with no modes at all — trading discoverability against a cleaner, more automatic experience.

- Without mode labels
- With mode labels — the direction we took
- Universal, without modes

## First run experience

Because the camera was not a familiar entry point to search, the first run had to educate users about a new way to search and earn camera access in the same breath.

![The first run introducing visual search.](https://varunv.com/media/2018/visual-search/section-9.webp)

![An example skill shown before permission is asked for.](https://varunv.com/media/2018/visual-search/section-10.webp)

![The camera access request, explained in context.](https://varunv.com/media/2018/visual-search/section-11.webp)

![The camera live for the first time after access is granted.](https://varunv.com/media/2018/visual-search/section-12.webp)

## Dogfood feedback and challenges at this step

1. Early ML models gave highly inaccurate and poor results.
2. Mismatch in user expectations versus the results.
3. Lack of finesse in the image capturing implementation.

Continuous iterations and improvements — in the UI, in the flow and at the backend — led to a solution where we as a team felt confident to release. Sensing, hold still, move closer and error states each told the user what the camera needed.

- Sensing
- Keep camera still for a moment
- Move closer
- Error

## Evolved live detection UI in V2

Intent cues appear as the camera scans the scene and objects, naming what it sees — plants, planters, jeans, jackets — before the user commits to a search.

![The camera naming several plants and planters in one scene.](https://varunv.com/media/2018/visual-search/section-13.webp)

![The camera naming the jeans and jacket a person is wearing.](https://varunv.com/media/2018/visual-search/section-14.webp)

## Shop the look

![An outfit in camera view, with shoppable points marked.](https://varunv.com/media/2018/visual-search/section-15.webp)

![The shop the look sheet, filtered by brand, price, size and colour.](https://varunv.com/media/2018/visual-search/section-16.webp)

![Matching products listed with prices and offers.](https://varunv.com/media/2018/visual-search/section-17.webp)

---

## Identify landmarks

![A landmark held in camera view and labelled.](https://varunv.com/media/2018/visual-search/section-18.webp)

![The landmark named, with a short description.](https://varunv.com/media/2018/visual-search/section-19.webp)

![Further reading and related places for the landmark.](https://varunv.com/media/2018/visual-search/section-20.webp)

---

## Identify paintings

![A framed painting detected in camera view.](https://varunv.com/media/2018/visual-search/section-21.webp)

![The painting named alongside its artist.](https://varunv.com/media/2018/visual-search/section-22.webp)

![Related works and context for the painting.](https://varunv.com/media/2018/visual-search/section-23.webp)

---

## Identify text and translation

![Printed text picked out of the camera view.](https://varunv.com/media/2018/visual-search/section-24.webp)

![The captured text translated in place.](https://varunv.com/media/2018/visual-search/section-25.webp)

![The translation offered with actions to copy or share.](https://varunv.com/media/2018/visual-search/section-26.webp)

---

## Live Text

![Live text selected directly in the camera view.](https://varunv.com/media/2018/visual-search/section-27.webp)

![A detected phone number and address ready to act on.](https://varunv.com/media/2018/visual-search/section-28.webp)

![The recognised text handed off to another app.](https://varunv.com/media/2018/visual-search/section-29.webp)

## Scaling the experience to Microsoft products

- Bing Search
- Microsoft Start
- Microsoft Edge

## Skill platform — 1st party skills developed by team

Find chairs like these. Find similar heels. Which tulips are these? Know about landmarks. Identify dog breeds. What painting is that? What's that dish? Celeb lookalike? Identified a lot of use cases with research.

- Find similar heels.
- Know about landmarks.
- What painting is that?
- What's that dish?
- DIY fix.
- Objects.
- Food, recipes and nutrition.
- Receipts.
- Colour.
- Antiques.
- Identify dog breeds.
- Celeb lookalike.
- Find chairs like these.
- Health.
- Cosmetics.
- Fix devices.
- Wines.
- Which tulips are these?

- DIY fix — “I want to know how to repair this section of concrete.” Today: go to Home Depot and ask, or search YouTube videos.

- Fix devices — where do I get printer ink, where can I buy a new printer, how do I fix my printer?

- Wines — “I want to easily find out about the wines when looking for a new one.” Today: manually search each one, which is tedious, or ask the store rep.

- Food, recipes and nutrition — where can I find recipes like this, how many calories, how do I record them for my diet?

- Cosmetics — “I want to easily find out about the cosmetics.” Today: manually search or ask the store rep.

- Health — where can I find bike workouts for my trainer, will the weather be good enough to ride outdoors, where can I get tubes to repair a flat tire?

- Objects — “I admired the coffee cup I was served and wondered where I might get something similar.”

- Receipts — help me call this business about my purchase, save this for a business expense, find the closest store, and see what rebates I can get.

- Antiques — “I admired an antique sculpture in a park and wondered what the history behind it is.”

- Colour — help me find objects that go well with this colour, and capture the exact hue of an object.

## Version 2. Taking visual search to the next level with scene intelligence and more comprehensive skills

- Understanding an image at a scene level, not just specific objects.
- Providing the ability to hand off to supporting apps that could complete the task at hand.
- Third party skills integration.

![A whole scene read at once rather than a single object.](https://varunv.com/media/2018/visual-search/section-30.webp)

![Several objects in the scene offered together.](https://varunv.com/media/2018/visual-search/section-31.webp)

![A skill picked up from the recognised scene.](https://varunv.com/media/2018/visual-search/section-32.webp)

![The skill carrying the task forward in place.](https://varunv.com/media/2018/visual-search/section-33.webp)

![A recognised task handed off to a supporting app.](https://varunv.com/media/2018/visual-search/section-34.webp)

![A third party skill surfaced inside the camera.](https://varunv.com/media/2018/visual-search/section-35.webp)

![The third party skill completing the task.](https://varunv.com/media/2018/visual-search/section-36.webp)

![The result returned to the camera view.](https://varunv.com/media/2018/visual-search/section-37.webp)

### Outcome & impact

- Positive movement in app retention.
- 20M monthly active users (MAU).
- Over time built comprehensive skills which were utilized elsewhere in the search ecosystem.

### Challenges and limitations

- The camera is not typically considered an entry point to search scenarios. How do we reduce friction in going from image to knowledge?
- How can we make visual search more top of mind, when Microsoft's presence on mobile was limited? This was more suited to being integrated into an OS.
- Engineering push back and challenges were immense. The team wanted to focus more on image based search rather than live detection.

### Direction for the future

- Make it reliable — elevate visual search as a more delightful and reliable search input.
- Ensure clarity in UX — reduce clutter and look to simplify complex interaction patterns.
- Stay grounded — an experience based on and tested with real world scenarios.
- Continually improve object detection — get better at understanding the nuances of the object.

---

Source: https://varunv.com/work/visual-search