Input type — Pictures
AI models for pictures
Twenty-four services that read a still image and return structured output — from a scanned form to a crowded square. One image over an API, or a whole archive in a batch, on the same dedicated cluster.
The right model for the job
Our AI advisors may assist you to select the best model or pipeline for the process or offer the most efficient hardware to perform the process.
Of course, you can bring your own model to run on our hardware.
In case you like to use your on-premise servers cluster or a certain cloud we may provision the needed services and build the pipeline.
AI services for pictures
Our service catalogue include vast selection of services.
Part of the services based on stable AI models and part of the services based on our native code.
Each service can use several models and methods. After a short discussion we can decide on the right pipeline for the process.
{{ filterNote }}
Documents and forms
- Handwriting recognition
- Handwritten notes, ledgers and filled fields read into text.
- Invoice / receipt extraction
- Supplier, dates, line items and totals as fields, not free text.
- Form extraction
- Key-value pairs pulled from structured and semi-structured forms.
- Table extraction
- Printed tables recovered as rows and columns, ready for a spreadsheet.
- Document classification
- Each page sorted into your own document types.
- Document visual question answering
- Ask a question of a page and get the answer with the region it came from.
- Document redaction
- Names, numbers and chosen regions masked before a document is shared.
Faces and people
- Face recognition
- Faces matched against a face database you hold and control.
- Face verification (1:1)
- Two images compared to confirm they show the same person.
- Face detection / face count
- Faces located and counted, with no identity involved.
- Crowd counter
- People counted in dense scenes where bodies overlap.
- Pose estimation (human)
- Body keypoints per person, for posture, gesture and activity work.
Objects and scenes
- General object detection
- Common objects located and labelled with a box and a confidence.
- Open-vocabulary object detection
- Describe what to find in words and it is found, without retraining.
- Image segmentation
- Pixel-level masks per region, for area and coverage measurement.
- Image classification
- A whole image tagged against your own label set.
- Visual similarity / duplicate image detection
- Near-identical and re-edited copies found across a large library.
- Industrial defect / anomaly detection
- Departures from a known-good reference flagged on the line.
Description and image quality
- Image-to-text captioning
- A written description of the scene, for search, archives and alt text.
- Image upscaling / super-resolution
- Resolution raised on small or compressed source images.
- Face restoration
- Blurred and degraded faces recovered from low-quality frames.
Hidden data and provenance
- Image steganography — hide text in photo
- A message carried inside a photograph, invisible to the eye.
- Image steganography — hide small binary data in photo
- A small payload such as a key or signature embedded in the pixels.
- Digital watermark / provenance embedding
- Origin and ownership marked into the image so it survives re-encoding.
How to provide us pictures?
We support any OS / Encoding and format for images.
You can upload your images with SCP (behind VPN) to a URL we provide.
You can provide us URL, MQ or any API and we will pull the images directly from your server.
You can send a NAS or a DAS to our location.
You can use our AI services with an API.
How to get the results?
We can provide the results with SCP (behind VPN) directly to your server, to use your API for results, send you SCP (behind VPN) details to download the results or sent you back the NAS/DAS you provided width the results.
Other input types
- Documents Extraction, classification, redaction, question answering. AI models for Documents →
- Video files Tracking, transcription, subtitles, summarization. AI models for Video files →
- Audio files Transcription, diarization, separation, summarization. AI models for Audio files →
- Live feed Counting, tracking, live transcription, PPE and queues. AI models for Live feed →
- Text files Summarisation, extraction, translation, PII detection. AI models for Text files →