Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking for recommendations on the stack used to build this, expecting to frown at them still recommending tesseract but I found that they all recommend apple's vision framework. I went back through old models and it seems like the last model to recommend Tesseract for a project like this was gpt-4.1.
Can you copywrite things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
Since this is for the mac you really should be using apple's vision framework for OCR. It smokes tesseract in both speed and accuracy.
Edit: I'm curious which LLM was used to generate the code. I fed the title of your post to claude/deepseek/qwen/codex asking for recommendations on the stack used to build this, expecting to frown at them still recommending tesseract but I found that they all recommend apple's vision framework. I went back through old models and it seems like the last model to recommend Tesseract for a project like this was gpt-4.1.
Slightly offtopic, but made me wonder.
Can you copywrite things like this now that LLMs exist? I mean, up until now if a small startup has a great idea they will get bought out by big tech which will integrate (or kill) their tech. But now with LLMs can the likes of OpenAI just tell their model to make something that works similar to X (such as this project) and then get round copying laws and negate being behind the curve?
copywriting is the art of writing copy for products/marketing maybe you were thinking of sherlocking [1]
1. https://news.ycombinator.com/item?id=34080326
I sure the intended word was copyright[0], as in to protect against getting sherlocked.
[0]https://en.wikipedia.org/wiki/Copyright
Why chose CLIP to do this. Have you tried small VLMs like Qwen-VL? I believe those models have video encoders can better perform at this scenario.
Having built something similar with CLIP on an M1, frame sampling rate is the whole ballgame. One frame a second on 12k videos is days, keyframes only got me to an overnight run.
Maybe you need a minimal downscale version as well, I heard is very common technique in the video editing world.
Based on my experience, sampling rate can be tricky if what you are looking for lasted less than interval period.
Proxies. You transcode proxies from the original media, edit off those, then you use OM for the final render. NLE’s usually let you flip between them.
Have you tried scene detection? I would guess camera cuts are even less frequent than keyframes.
> 12k videos ?
you are pirating first-release movies for commercial purposes?
I think 12K is quantity not resolution.
How well do you think this would work on stock photography on m1 mac with 32GB ram? For example I'd like to be able to search a folder of ~2k photos for houses with palm trees. Or find photos of kitchens, or find photos of desert southwest landscapes.
possibly off-topic, but for anyone interested in this on a more cross-platform / holistic basis, Immich does this
(& by "this" I mean an approximate AI search for photos & videos - I can't account for the "every frame", nor for the comparative search quality)
I like the entire premise, the one thing stopping me from trying this is not knowing the time scales that I will need to set my computer aside for the processing of large folders of video frames, or my photos library's videos, some 12,000 videos
would be lovely if picture embeddings were attached to the file by the camera but one can only dream of such futures
This is cool. Any way to search for People / faces / pets? Like on iOS?
How do you do that? What's the architecture? Can you guide on that?
Why is this a JS bloatware instead of native or Rust which is easier than ever now with LLM coding tools.