This is the coolest thing on HN that I've seen in a minute. The whole thing is a perfect blend of ideas with the end result of something that feels... magical. Honestly this is the highest inspiration to me as a builder as I just want to create magical experiences, even if they are just small oddities.
That's the e-ink monopoly for you. Though you could probably use a slightly smaller b&w eink display with a second hand raspberry 4 to shave a lot from that number.
I had not heard it before (or if I did I did not interpret it correctly, eg if one told me "see you in a minute" I would interpret it as "see you in a bit".
But there is this podcast discussing it in 2021, and it comes from black community slang from 70s (which is where most slang I encounter comes from). And these terms take a while to catch up usually, but it seems it was already circulating more broadly since 2000s.
It's been around in slang for a few years with a few variants, e.g. "I haven't seen her in a hot minute." I think it's just filtered into wider cultural vernacular.
Are there any LLMs being widely used for audio classification? I know VLMs are being used a lot in image stuff.
It always seems kind of silly to me to throw everything at an LLM. I know they’re huge and can automatically handle a huge number of tasks but something in me finds it wasteful when we could be creating easily trainable, cheap to run bespoke models for a lot of stuff
There are LLMs that support audio input, similar to those with vision support.
From my testing of open weights LLMs with audio support, they basically are only trained to recognize audio as an alternative to text input, they treat audio as basically equivalent to a transcript, and can't recognize or distinguish things like music, accents, background sounds, etc.
So they're only really good for transcribing or summarizing or using audio input in place of text input for prompts, but not anything that requires distinguishing any information about the audio that would not be present in a transcript.
It can be tempting to try to use an LLM for a variety of tasks; kind of the whole thing about an LLM is that you don't have to do a separate complex training run for every task, but can just provide instructions in natural language. But it only works as far as what the training data covers, if the training basically always treated audio and a text transcript as equivalent, the model has nothing causing it to learn other relevant features of the audio. If there's enough bird call identification in the training data of an LLM, it might be able to do that, but I think multimodal training data tends to be much more limited than the text training corpus
Have you had any luck fine tuning one with musical data for classification or music-aware QA? I've been hacking on https://trebel.la/ which I would like to be a music practice companion, and the biggest missing feature is actually useful audio-based feedback pipeline.
My current approach, not yet validated, is trying to generate training data from masterclass recordings on Youtube, and then fine tuning MOSS-audio on a bunch of those. But I'm interested if there are better models, or large training sets I don't know about.
I'm not selling anything; I'm developing it live and there is a landing page, but there's no payment hooked up since like you said, it's not functional yet.
The harness that connects to a chatbot, API or voice interaction is the place to route requests to different systems. If you remember the early days of ChatGPT it explicitly said it was routing image generation to Dall-E after embellishing your request itself first.
Determining which tool to use should be a lightweight operation but I’m not expert enough to understand exactly how much lighter than a full LLM call just to recognize it needs a different tool or model.
For sure, I just know it’s tempting given the power of large transformers to throw things at an existing model.
For instance, OCR is something that can be done locally with no access to a GPU but people (including me) still often use cloud hosted multi-modal large language models for it.
wav2vec or similar approaches are used a lot these days, which is basically BERT with audio inputs. Whether that counts as an LLM or not, I am not sure. It's a transformer architecture in any case.
People will often reach for "easily trainable, cheap" solutions when they can; the reason people reach for Transformers and LLMs is because when you throw more data at them, they get better.
Note that while the underlying birdnet-go project started as BirdNET only, it can now use Google Perch v2, BattyBirdNET (for bats!) and other models in the future. It's a really cool project!
Why would anyone assume this uses an LLM? It classifies bird sounds, not human language. I don't mean this as an attack, I'm genuinely curious! This seemed obvious to me, and I want to know what line of thinking might lead one to believe that an LLM is the better (or more likely) tool for this job over a purpose-built classifier.
I wanted to say that seems like a stretch but then I often find myself visualizing my best guess at the appearance, including species, of an unknown dog. It seems like size of dog and pitch of bark are negatively correlated
e-ink is so much fun, especially when combined with ESP32 or BTLE boards. I currently have 4 around my house displaying book quotes that I've highlighted in KOReader and they bring me joy each time I see them. They are simple and do one thing very well.
I just got and setup a BTLE e-ink driver the other day and according to my calculations it will last _years_ on a single charge (2000mAh) even with multiple refreshes per day (crazy compared to the Wifi ones). Year+ lifetime completely changes the calculus for where to put these IMHO.
I'm considering designing a multiple-screen "art" piece to hang on my wall and playing around with various ideas (on both content and design).
LLMs + 3D Printer + e-ink is my new favorite hobby.
Here are the parts I bought, I'm still getting it all setup (I have my pictures displaying, working on designing the 3D printed case now, and testing it all) but according to OpenDisplay's calculator [0] I can expect 3.4 years if I refresh the screen once every 4 hours (and currently I'm doing closer to once a day to mimic my ESP32s which need deep sleep and once-a-day refresh to get even month/months of battery on a charge).
The image rendered to the screen is coming from a little service I wrote that renders a book quote it pulls from BookOrbit which syncs with KOReader (annotations, progress, etc). My plan is to set up one of these one each bookshelf/series rotating quotes I highlighted from the series.
Here are 2 ESP32-based screens I have, the one on the left is a an all-in-one [1] and the one on the right I printed the case from the TRMNL DIY kit from SeeedStudio [2]: https://cs.joshstrange.com/knk4Ns2h
Right now only the BTLE one is running OpenDisplay (the other two are semi-managed by HomeAssistant ESPHome) and all the devices are in HA for controlling them, though the ESP32's all deep-sleep for saving battery so I can can't "push" stuff to them like I can the BTLE one.
If I want a wifi e-ink connected display - curious what would you recommend.
I want an e-ink display - that I can use as a sign/display board with wifi connectivity and remote administration.
If you want a simple e-ink wifi screen I'd point you at the reTerminal e1001 [0] if that size works for you. It's the easiest to get started with. You can either keep it on power or use the deep sleep to get more from the battery. It really comes down to how often you want to update it.
The BTLE boards have just crazy higher battery time with no need for deep sleep which is why I'm moving in that direction.
~7.5" is the size I've been buying so far, the larger e-ink screens get pricey but I can't wait for a calendar-sized one to come down in price such that I can frame it and hang it on the wall. Lastly I'd avoid the color e-ink screens unless you know you need color, won't need to refresh it often, and it's not in your eye line. Color refreshes take way longer (~30s in my experience for a 5-6 color screen) and the flashing it has to do is obnoxious. Monocolor displays take <2s normally and a single flash. I got a color screen first because "why not?" and I got color, but the refresh was too annoying so now it refreshes once a day in a place that I'm not when it refreshes so I don't have to see it.
Thanks for speaking to the battery question; I'm very interested in how to run this kind of thing (the display portion at least) cordlessly too, especially when lipo pouches in that range are so dirt cheap.
I might be misunderstanding your comment but the display is run/powered by the board so all you need to do is provide power to the board. For e-ink it uses no power to "hold" an image, only to write/refresh (sorry if I'm telling you something you already know!).
So reiterate, with a 2000mAh lipo battery I should be able to drive the board _and_ the screen for over 3 years if I only refresh it at most once every 4 hours (with "push" capabilities, I don't have to deep-sleep the board to get that battery life). At least, that's what the calculator says, I _just_ got my hardware to play with and so I can't speak from experience. Unfortunately my board is not charging the battery so I need to get a replacement which will probably take a week or two to get here.
Yeah no we're aligned; I have some raspis that I'm interested in doing long term battery stuff with (remote sensing and the like) and it definitely does seem like connectivity is the huge battery killer, like even if you wake up only periodically to take a reading or update an epaper display, the cost in battery just to negotiate a wifi connection and do tcpip and http things is awful.
And then you're stuck with a device that can only be communicated with when it chooses to wake up, whereas with BLE it seems you can listen for a remote OTA wake much more cheaply.
Going to google/search and stuff when I get home, but for my current living space it would be awesome if I could get a wireless mic to place outside and have this run on my Beelink SER8, and then over local network sync to my Apple TV and display the images kinda like a screensaver?
Awesome! It's really cool that it uses public domain images as the inspiration, too. And even cooler that you link to the places where you got those public domain images.
My mistake for calling it "copy", should have said fork maybe (?)
They are awful close to still require a reference to the original project imo. They are achieving the same output (identify birds and display them in a frame display all hosted in a raspberry pi) with that twist to how to generate the art.
Author here: I'll take any feedback and critisism. I haven't based any of my work directly on his project, but will be referencing it as a similar one in the readme, as well as e.g. https://github.com/veteranbv/inky-bird-frame.
From the link: "Half the point of this project is showing off some amazing public-domain natural-history illustrations. Over 800 cut-outs covering more than 400 species, every one taken from a real plate and hand-curated for this project (no art is AI-generated, though some has been retouched with AI)."
Idk what "retouched with AI" means. Does it mean they used a segmentation model to cut out the shapes, or were they shoved through an imagegen and completely recreated?
I agree I would expect a reference. Another difference between this an AvianVisitors is the image sources. AvianVisitors is all AI generated with a consistent prompt, but OP's is using chiefly public domain images with retouching by AI as needed.
I understand OP is continuing the work and improving it as he sees fit, but that in my mind is a fork and should aknowledge the initial work that sparked it.
Maybe i'm wrong tho and this is completely parallel work! Waiting for OPs reply
Hi, yeah I'm not basing any of my coding or feature-work on that project. My biggest source of inspiration is actually just a paper poster I have hanging on my wall https://www.axelthorenfeldt.com/news/wwf-verdens-naturfonds-.... I wanted to make a version of that, showing the actual birds in my garden.
A fork implies that it builds upon code from another repository. I wouldn't use the word fork to mean "inspired by". I think it takes away from the effort that the author has put in. Anyone can create a fork of a project with the click of a button.
Original author here. Fair point, this project was one of my inspirations, but I chose to take a different approach. This uses public domain natural history art and builds on top of BirdNET-Go which is a popular way of self-hosting bird-detections. (This is really a BirdNET-Go companion, as it cannot run without).
That's fair, and I totally get wanting to give it your own spin.
I still think it would be good to acknowledge the original project that sparked the idea, especially now that you've mentioned it was one of the inspirations.
Definitely keep going with yours and evolve it further. That's how good projects grow. But for a healthy open source community, I think we should make a point of crediting the people whose work inspires us. Otherwise, over time, it can discourage people from sharing their projects publicly in the first place.
Contrary to the Great Man version of how we mostly talk about history, scientific discoveries, technological breakthroughs and such have historically often occurred multiple times in multiple places often in close proximity time-wise but otherwise completely unrelated to each other. This of course makes sense if you consider that history seems to be primarily driven by material conditions (e.g. steam engines were known to Ancient Greece but their usefulness for industry simply hadn't occurred to them because they lacked the prerequisites to make them useful and neither economic efficiency nor access to cheap labor were significant enough concerns - plus there was no patent system of course).
I'll leave it up to you to decide whether one should expect the demographic of this website to be especially invested in the Great Man narrative but trying to convince them that two people can independently have similar-sounding ideas involving similar-but-different technology is probably an uphill battle. Just be glad you're not composing music for a living because there's literally an entire industry built around the idea that it's impossible for two people to come up with similar melodies without knowingly or unknowingly copying eachother and that it's very important that one of them (or rather the company they had to sign a deal with in order to make any money off their music at all) 100% owns that exact melody and should be paid by anyone who wants to use it.
Next time you have an idea you should just file a patent tbh - the insistence on prior art invalidating the work seems to be quite similar but at least if you're granted a patent you can shake down "copycats" for money.
> Contrary to the Great Man version of how we mostly talk about history, scientific discoveries, technological breakthroughs and such have historically often occurred multiple times in multiple places often in close proximity time-wise but otherwise completely unrelated to each other.
My favourite example of this is calculus, followed closely by heavier-than-air flight.
I think there is ample evidence that the Wright brothers were (unlike many other inventors) uniquely suited to solving the problem of heavier-than-air flight, as they solved several open problems that nobody was making progress on (the main ones being control authority, propeller designs, the scientific basis for wing designs, and the ability to test designs and practice without dying) -- they even discovered and solved previously-unknown problems (such as adverse yaw, something their contemporaries should've independently discovered if they were on the right track).
A full rundown would be too long but [1] is a pretty good video going through the history of flight and how most of the other candidates for "first flyers" people bring up don't count and how it's difficult to argue that any of their contemporaries were even close to the Wrights.
When I first heard of it (through a community member in discord), I wanted to see if I could replicate it on FrameOS (my software for running anything on an e-ink panel), and sure enough, I got quite close: https://scenes.frameos.net/s/bird-field-journal
This is obviously nowhere as complicated as OP's project, but I thought I'd just share that your project inspired me to build something similar too.
For those of you that have a Samsung frame tv and no spare e ink screens:
Inspired by AvianVisitors [1] I made a similar project [2] with an android app for bird detection using Perch, and then displays the birds detected during the last 24 hours with nice art on a samsung frame tv in "art mode".
If you dont want to repurpose an old android device, my app also supports bird detections from birdnet-go and birdweather. It should be plug and play with this project.
That's indeed a cool project and it's this kind of stuff that makes HN a unique corner in this digital world. I can classify this project with [1](Ranked #3 as the most upvoted Show HN of all time), [2](#4) and [3](#6) as non pure-software products that happen to also integrates hardware as well.
The only thing I wish to be added is a short video showing the actual frame in action to express motion and sound so I can better feel it. While the provided image (hero.jpg) gives me a sense of the product, I just can't smell it yet. I visited https://fugleramme.arnegiacomo.dev/ hoping for a live feed but it seems just a static image.
Kudos for the author and I hope we can see the next iteration!
I've been leveraging Birdnet-2.4 as well, but also ultrasonic-pass and activity-v1 to do broader detection with an ecological targeted microphone. This means I can capture not just bird sounds but also bats and some insects, and acoustic events such as jets and cars.
Applying heterodyne transformations to the bat recordings makes them audible which is fascinating, especially when you catch them in a feeding frenzy.
Similar intent with an indoor display, although nowhere near as beautiful as OPs eink display here. I'm more aiming for a real-time ticker, but this design is wonderful. I'm quite inspired by it!
I am working on a voice-chat program and I implemented RNNoise this week. Most of the noise cancellation tools filter for just human voice. so bird noises that you want, may be gone when you are trying to delete the wind noise or something. And without any noise cancellation, It can be very hard to people who are living next to road or very noisy place.
I haven't tried with BirdNet, but Merlin doesn't seem to ever false-positive on chipmunks or squirrels (there's been a suggestion that it should explicitly recognize them, and maybe some frogs too, since they're "loud thing in the same register" and it frustrates people that haven't figured it out yet...)
I wish E-ink displays were a bit more affordable. The 13 inch one is £229.50.
Does anyone know why they are still so expensive, at least at large sizes?
You have to distinguish the fake color e-ink that are used in mass consumer products which are actually black&white e-ink plus an LCD layer and are expensive, and the real color e-ink where each micro capsule embeds 3 to 5 different inks of different colors.
The latest are pretty complex devices, very slow to refresh (hence why they are not used in your e-reader) and also very expensive. But the rendering is way much nicer because it’s real ink that is used to render the colors.
I don't see LCD layer here. Passive color filters, B/W capsules (which reflects or doesn't reflect incoming light, depends on orientation), TFT (thin layer transistors, not LCD — liquid crystal display!) as control to orient capsules.
Yes, it is not 4 color capsules as in Spectra, but where is LCD?
LCD what is caught my attention (and surprise) in your commentary.
I am using my 13 inch inky display for my own birdnet - I also posted this in another thread last week https://i.imgur.com/5XbM6bb.png
I also bought 3 more 4 inch ones. Here is the first deployment next to the fish tank - https://i.imgur.com/nj8nKrS.png - I like this because I can stuff an rtl sdr in that box too, so it picks up nearby radio sensors in addition to BT. Real nice project screen!
I usually only refresh the image on the screen every few hours. I love how the image stays even with no power.
Any easy way to get started on your own version of this without needing a mic would be to pull birds seen nearby using the eBird API and pairing scientific names. Won't be as accurate as a backyard mic, but will give you a direction to start.
Also if you have old e-ink device laying around, you can likely setup BYOD with TRMNL (or just buy one https://trmnl.com/) and send the bird images as a private plugin
Thanks for sharing. Didn’t know about this project. Have you used it yourself? I’ve only experimented with Scrypted for fun, but Frigate sounds like I could operate birdsong analysis without a separate BirdNET container.
I have Bird Classification enabled in the settings, but so far haven't seen any labelled (it's only been a few days). Similarly with Face Recognition, I've only seen false positives. It does reliably detect people though (300 tracked so far). I'll have to experiment with settings such as confidence score.
What would be cool is if you can put AI to identify individual birds. They all have different calls/voices, though largely indistinguishable to human ears. And you can name them and see when they come for a visit.
This is an excellent execution and refreshing to see the thought process behind it. Kudos. the ideas such as these, gonna keep us swim sane and specifically extending far beyond the 9-5 is no longer sideProject category but refreshingly needed one.
I just happen to have an extra Pi and an e-ink display sitting around and I LOVE this idea. Going to try it on my screened in porch - we typically get 10-20 Merlin detections on a decent day so this could be fun!
On a side note. How do people deal with all the ugly mess birds cause in such settings? I'd love to have such thing but seeing all the white stuff on glass door does not really feel fun to me overall.
What do you mean specifically? These birdsong analyzers are sensitive enough that they do not rely on setting up a feeder. It’s obviously dependent on where you live, but in the Nordics, living outside city centres is good enough for both background noise and birdsong from nearby trees/parks/forests. Privacy matters aside, setting a mic up on a balcony or window is enough. Another solution is feeding data from a cheap security camera with a mic through wifi. This way the camera can stay outside in different weather, it’s not super close to you actual living area and splitting audio from the RTSP stream should be relatively easy to be analyzed by BirdNET.
Well I specifically meant having a bird feeder on a window or glass door for a chance to watch them live closely. That's why I put it as a side note ;)
You gave me few ideas for the analyzer setup too - thanks!
Yeah that’s definitely going to be messy. Based on your region, there should hopefully be generic security cams that can even work with a batter and solar. That combined with wifi if it can reach your house can hopefully steer the blast zone way further :D There’s a few TP-Link cameras on discount around 100€ that have 2k res. Feels a lot easier this way rather than taping together a weatherproof box for various components like the Pi and a mic. So far I’ve only used a plug-powered one since I have a IP-rated socket outside and the TP-link cam has a magnet for my wall. Hope you find a nice solution!
Sounds like you've got birds crashing into the door? That's often something that can be improved (with stickers or hanging decorations) that make the glass more perceptible to the birds. (Another thing that sometimes helps is having bushes or shrubs, or even potted plants, in sight of the feeder - most small birds will use a "staging area" and fly somewhere nearby first, "check out the scene", and then fly to the feeder from there; it doesn't really control the activity, but it adds some options.)
This is amazing. I want to build this. Curious about the authors “full garden” link, surely he doesn’t get that many birds in such a small space of time in his garden, unless he lives next to a park I suppose?
Author here. A mic listens to the sounds in my garden, BirdNET-Go (an open-source BirdNET classifier) detects birds by sound, and the e-ink frame draws a real 1800s natural-history illustration of each one as it hears them. Here’s a live web-demo from my garden in Bergen, Norway: https://fugleramme.arnegiacomo.dev
A few things that might be interesting:
- The screen only redraws when the set of birds changes, and dithers the collage down to six colours for the e-ink display.
- Bigger birds sit toward the centre, scaled by real body mass.
- 800+ cutouts across 400+ species, each cut from a real public domain plate. All art is historic and human-made. Coverage is currently best for the Nordics, Britain and Germany (but other parts of the world are in the works).
- Runs fully local on a Raspberry Pi or your homelab (yes, even the classifier runs great on a RPI)
The e-ink panel is optional and it can run web-only in a container against a BirdNET-Go you already have.
If I were to make one of these as a Christmas present for somebody in Texas, and I wanted to contribute work toward good coverage for the area, how big of a time commitment would that be? It sounds like the kind of hobby work I could get into for a few months.
Less than you'd think, because you likely wouldn't need full "Texas coverage", you mostly need the birds that actually show up. BirdNET filters by location, you'll likely need a couple of dozen. Over the summer I've had about 40 visitors where I live.
Right now the coverage of North America is relatively sparse, but the interest has been high and I'm hoping for some contributions.
Edit: For North America the most obvious source is Audubon's Birds of America, which is public domain and should provide a lot of coverage. This collection includes some beautiful plates.
Do you have your image source available somewhere? I'm not going to build this same thing, but I'd love to add some prints of birds to my shelves and, not knowing much about birds, it would help to have a set of pictures I could look through and go, "Ooh! This one!" and learn about them. It'd be a good jumping on point!
I'm actually currently working on implementing Iberian species, and added some yesterday. There's another user also contributing with some beautiful Spanish birds. Feel free to contribute if you have some cool visitors in your garden!
I can't help wonder what this would look like for humans. Imagine, set one of these up at a party and it has the means to identify anyone on the invite list. When it hears the person it displays a photo, brief bio, and relationship status.
We could also leave fake images aside, go outside and try to watch the birds ourselves. Maybe even let the camera inside, just looking with our own eyes. Real nature. No intermediate.
> Half the point of this project is showing off some amazing public-domain natural-history illustrations. Over 800 cut-outs covering more than 400 species, every one taken from a real plate and hand-curated for this project (no art is AI-generated, though some has been retouched with AI).
You can in fact both have an interesting bird-based art display in your home and also go outside and watch birds yourself. Neither precludes the other.
While I love this idea generally, the scope of the use of generative AI for artwork is too vague for my comfort.
The README.md mentions that "no art is AI-generated, though some has been retouched with AI." Superficially, that sounds mild. Some of the other closely-related projects, however, use different verbiage, or none at all.
For those of us with ethical concerns around generative AI application to artwork, much more detail may be needed before we could decide if this was appropriate for us.
Nothing at runtime uses AI either, at least in the sense of LLMs or diffusion. The collage is Pillow and numpy, and detection is BirdNET, which is a classifier, not a generative model.
However I have used LLMs for code-related work, and for writing scripts for programatically editing the images, like cutting, contrast and colour corrections.
Finally, we solved the problem of converting bird sounds into their pictures without involving human cognition. The great problem and the need had to wait for the great, ultra expensive, data-center-driven AI technology to appear on the horizon. AI is such a good technology that continue to solve many such great problems that humans desperately needed for their survival and progress. All at the a minimal cost of some effect on the climate, which may only slightly quicken the human extinction.
Puzzlingly, no animal species requires AI, which is a mystery that need to be solved by AI.
Hmm. The project uses BirdNET-Go, which does inference locally on a Raspberry Pi. Then it looks up the species against a set of already existing illustrations. So in this case it seems like the "ultra expensive, data-center-driven AI" is not really involved in the core loop of this project.
This is incorrect; the local model is a small dense neural network called BirdNET that would not have been referred to "AI" when it was published in 2021. The model outputs a bird species probability distribution and the web application displays an existing image file of the most probable bird on the screen. This is a lovely example of simple, offline, and fun project using a straightforward machine learning model.
Ok, so why didn't this small dense model come up, say, 10 years back? Oh it needed all the AI evolution and concepts that were powered and evolved inside the data centers. But we still want to claim that it has nothing to do with the Big AI.
What a shame that you use that excuse to offset all other evils of AI, without even having a hint of how many patients were saved by AI. Just like how oil salesmen and nukes makers say that they solve some great problem of the world.
Incorrect. As others have pointed out, BirdNet is a “traditional” neural network. The amount of compute needed to train something like this is many orders of magnitude less than an LLM. Something like this could be trained on local hardware with enough juice, or by renting a handful of GPU’s
Probably would not show much in the U.S.
Off-topic, but what is up with the increased use of the phrase "in a minute", presumably to mean "in a long time", lately?
I've only started encountering it in the past year.
Did it get popularized by some celebrity, tv show, influencers, etc?
But there is this podcast discussing it in 2021, and it comes from black community slang from 70s (which is where most slang I encounter comes from). And these terms take a while to catch up usually, but it seems it was already circulating more broadly since 2000s.
https://waywordradio.org/its-been-a-minute/
https://trends.google.com/explore?q=I%27ve%20seen%20in%20a%2...
They took a well-defined unit of time, which is relatively short, and made it mean “some unknown but very long period of time”.
So frustrating. /oldmanyellsatcloud
https://doi.org/10.1016/j.ecoinf.2021.101236
It always seems kind of silly to me to throw everything at an LLM. I know they’re huge and can automatically handle a huge number of tasks but something in me finds it wasteful when we could be creating easily trainable, cheap to run bespoke models for a lot of stuff
From my testing of open weights LLMs with audio support, they basically are only trained to recognize audio as an alternative to text input, they treat audio as basically equivalent to a transcript, and can't recognize or distinguish things like music, accents, background sounds, etc.
So they're only really good for transcribing or summarizing or using audio input in place of text input for prompts, but not anything that requires distinguishing any information about the audio that would not be present in a transcript.
It can be tempting to try to use an LLM for a variety of tasks; kind of the whole thing about an LLM is that you don't have to do a separate complex training run for every task, but can just provide instructions in natural language. But it only works as far as what the training data covers, if the training basically always treated audio and a text transcript as equivalent, the model has nothing causing it to learn other relevant features of the audio. If there's enough bird call identification in the training data of an LLM, it might be able to do that, but I think multimodal training data tends to be much more limited than the text training corpus
My current approach, not yet validated, is trying to generate training data from masterclass recordings on Youtube, and then fine tuning MOSS-audio on a bunch of those. But I'm interested if there are better models, or large training sets I don't know about.
Determining which tool to use should be a lightweight operation but I’m not expert enough to understand exactly how much lighter than a full LLM call just to recognize it needs a different tool or model.
For instance, OCR is something that can be done locally with no access to a GPU but people (including me) still often use cloud hosted multi-modal large language models for it.
People will often reach for "easily trainable, cheap" solutions when they can; the reason people reach for Transformers and LLMs is because when you throw more data at them, they get better.
I just got and setup a BTLE e-ink driver the other day and according to my calculations it will last _years_ on a single charge (2000mAh) even with multiple refreshes per day (crazy compared to the Wifi ones). Year+ lifetime completely changes the calculus for where to put these IMHO.
I'm considering designing a multiple-screen "art" piece to hang on my wall and playing around with various ideas (on both content and design).
LLMs + 3D Printer + e-ink is my new favorite hobby.
Uses a native C driver instead of python so you can run the display off a battery as it's more efficient.
Screen: 7.5" Monochrome eInk / ePaper Display with 800x480 Pixels (https://www.seeedstudio.com/7-5-Monochrome-ePaper-Display-wi...)
BTLE board: XIAO ePaper Display Board(nRF52840) - EN05 (https://www.seeedstudio.com/XIAO-ePaper-Display-Board-nRF528...)
Battery: 3.7V 2000mAh - (https://www.amazon.com/dp/B0FR9GH966)
Picture of it all assembled (no case yet): https://cs.joshstrange.com/zJFvGGPB
The image rendered to the screen is coming from a little service I wrote that renders a book quote it pulls from BookOrbit which syncs with KOReader (annotations, progress, etc). My plan is to set up one of these one each bookshelf/series rotating quotes I highlighted from the series.
Here are 2 ESP32-based screens I have, the one on the left is a an all-in-one [1] and the one on the right I printed the case from the TRMNL DIY kit from SeeedStudio [2]: https://cs.joshstrange.com/knk4Ns2h
Right now only the BTLE one is running OpenDisplay (the other two are semi-managed by HomeAssistant ESPHome) and all the devices are in HA for controlling them, though the ESP32's all deep-sleep for saving battery so I can can't "push" stuff to them like I can the BTLE one.
[0] https://opendisplay.org/index.html#battery
[1] https://www.seeedstudio.com/reTerminal-E1001-p-6534.html
[2] https://www.seeedstudio.com/TRMNL-7-5-Inch-OG-DIY-Kit-p-6481...
The BTLE boards have just crazy higher battery time with no need for deep sleep which is why I'm moving in that direction.
~7.5" is the size I've been buying so far, the larger e-ink screens get pricey but I can't wait for a calendar-sized one to come down in price such that I can frame it and hang it on the wall. Lastly I'd avoid the color e-ink screens unless you know you need color, won't need to refresh it often, and it's not in your eye line. Color refreshes take way longer (~30s in my experience for a 5-6 color screen) and the flashing it has to do is obnoxious. Monocolor displays take <2s normally and a single flash. I got a color screen first because "why not?" and I got color, but the refresh was too annoying so now it refreshes once a day in a place that I'm not when it refreshes so I don't have to see it.
I also recommend you check out OpenDisplay for what you will run on the device: https://opendisplay.org/index.html
[0] https://www.seeedstudio.com/reTerminal-E1001-p-6534.html
So reiterate, with a 2000mAh lipo battery I should be able to drive the board _and_ the screen for over 3 years if I only refresh it at most once every 4 hours (with "push" capabilities, I don't have to deep-sleep the board to get that battery life). At least, that's what the calculator says, I _just_ got my hardware to play with and so I can't speak from experience. Unfortunately my board is not charging the battery so I need to get a replacement which will probably take a week or two to get here.
And then you're stuck with a device that can only be communicated with when it chooses to wake up, whereas with BLE it seems you can listen for a remote OTA wake much more cheaply.
Going to google/search and stuff when I get home, but for my current living space it would be awesome if I could get a wireless mic to place outside and have this run on my Beelink SER8, and then over local network sync to my Apple TV and display the images kinda like a screensaver?
It seems IP over Avian Carriers is finally within reach!
0. https://github.com/tphakala/birdnet-go
https://www.birds.cornell.edu/home/merlin/
https://theodore.net/store/avian-visitors/
Then have the screen look like a biologists notebook where he sketches the bird and writes a few notes about the time and its behavior.
https://github.com/Twarner491/AvianVisitors
https://theodore.net/projects/AvianVisitors/
Or maybe im missing to see is the same author? It went viral a few months ago
And just because 2 things do the same job doesn't make one a "copy" of the other.
They are awful close to still require a reference to the original project imo. They are achieving the same output (identify birds and display them in a frame display all hosted in a raspberry pi) with that twist to how to generate the art.
I understand OP is continuing the work and improving it as he sees fit, but that in my mind is a fork and should aknowledge the initial work that sparked it.
Maybe i'm wrong tho and this is completely parallel work! Waiting for OPs reply
Avian Visitors - https://news.ycombinator.com/item?id=48343424 - May 2026 (20 comments)
I've put a link to the earlier project in the toptext above.
I've seen other similar projects pop up like https://github.com/veteranbv/inky-bird-frame and https://github.com/adamoberley/HABirdDashboard/tree/HABirdDa..., each with their own spin.
I still think it would be good to acknowledge the original project that sparked the idea, especially now that you've mentioned it was one of the inspirations.
Definitely keep going with yours and evolve it further. That's how good projects grow. But for a healthy open source community, I think we should make a point of crediting the people whose work inspires us. Otherwise, over time, it can discourage people from sharing their projects publicly in the first place.
I'll leave it up to you to decide whether one should expect the demographic of this website to be especially invested in the Great Man narrative but trying to convince them that two people can independently have similar-sounding ideas involving similar-but-different technology is probably an uphill battle. Just be glad you're not composing music for a living because there's literally an entire industry built around the idea that it's impossible for two people to come up with similar melodies without knowingly or unknowingly copying eachother and that it's very important that one of them (or rather the company they had to sign a deal with in order to make any money off their music at all) 100% owns that exact melody and should be paid by anyone who wants to use it.
Next time you have an idea you should just file a patent tbh - the insistence on prior art invalidating the work seems to be quite similar but at least if you're granted a patent you can shake down "copycats" for money.
A full rundown would be too long but [1] is a pretty good video going through the history of flight and how most of the other candidates for "first flyers" people bring up don't count and how it's difficult to argue that any of their contemporaries were even close to the Wrights.
[1]: https://www.youtube.com/watch?v=EkpQAGQiv4Q
When I first heard of it (through a community member in discord), I wanted to see if I could replicate it on FrameOS (my software for running anything on an e-ink panel), and sure enough, I got quite close: https://scenes.frameos.net/s/bird-field-journal
This is obviously nowhere as complicated as OP's project, but I thought I'd just share that your project inspired me to build something similar too.
For those of you that have a Samsung frame tv and no spare e ink screens:
Inspired by AvianVisitors [1] I made a similar project [2] with an android app for bird detection using Perch, and then displays the birds detected during the last 24 hours with nice art on a samsung frame tv in "art mode".
If you dont want to repurpose an old android device, my app also supports bird detections from birdnet-go and birdweather. It should be plug and play with this project.
Have a look at [1] https://theodore.net/projects/AvianVisitors/ [2] https://github.com/simenf/birdframe
The only thing I wish to be added is a short video showing the actual frame in action to express motion and sound so I can better feel it. While the provided image (hero.jpg) gives me a sense of the product, I just can't smell it yet. I visited https://fugleramme.arnegiacomo.dev/ hoping for a live feed but it seems just a static image.
Kudos for the author and I hope we can see the next iteration!
_______________
1. I made an open-source laptop from scratch (https://news.ycombinator.com/item?id=42797260)
2. I replaced a $120k bowling center system with $1,600 in ESP32s (https://news.ycombinator.com/item?id=48968606)
3. A retro video game console I've been working on in my free time (https://news.ycombinator.com/item?id=19393279)
Applying heterodyne transformations to the bat recordings makes them audible which is fascinating, especially when you catch them in a feeding frenzy.
Similar intent with an indoor display, although nowhere near as beautiful as OPs eink display here. I'm more aiming for a real-time ticker, but this design is wonderful. I'm quite inspired by it!
I've been working on this in the background for a few months, if anyone is interested repo with screenshots is here: https://github.com/simonjgreen/OpenObservatory
I love HN
Genuinely giving the OP props
I wish E-ink displays were a bit more affordable. The 13 inch one is £229.50. Does anyone know why they are still so expensive, at least at large sizes?
The latest are pretty complex devices, very slow to refresh (hence why they are not used in your e-reader) and also very expensive. But the rendering is way much nicer because it’s real ink that is used to render the colors.
Rather opposite: I never seen consumer products with two-layer displays.
Vs Spectra (the expensive screen): https://www.eink.com/upload/2023_11_14/3_20231114082849vy2as...
Yes, it is not 4 color capsules as in Spectra, but where is LCD?
LCD what is caught my attention (and surprise) in your commentary.
But now I see difference, thank you.
Disclaimer: I am the CEO of Pimoroni :-)
I am using my 13 inch inky display for my own birdnet - I also posted this in another thread last week https://i.imgur.com/5XbM6bb.png
I also bought 3 more 4 inch ones. Here is the first deployment next to the fish tank - https://i.imgur.com/nj8nKrS.png - I like this because I can stuff an rtl sdr in that box too, so it picks up nearby radio sensors in addition to BT. Real nice project screen!
I usually only refresh the image on the screen every few hours. I love how the image stays even with no power.
See https://github.com/arnegiacomo/fugleramme/blob/main/assets/a... for all the sources
Also if you have old e-ink device laying around, you can likely setup BYOD with TRMNL (or just buy one https://trmnl.com/) and send the bird images as a private plugin
https://github.com/defl/hokku_epaper
On a side note. How do people deal with all the ugly mess birds cause in such settings? I'd love to have such thing but seeing all the white stuff on glass door does not really feel fun to me overall.
You gave me few ideas for the analyzer setup too - thanks!
Refreshes every hour
Battery lasts a few months
Esp32 and a colour eink panel, same as this one I think
Amazing thing, will replicate.
A few things that might be interesting:
- The screen only redraws when the set of birds changes, and dithers the collage down to six colours for the e-ink display.
- Bigger birds sit toward the centre, scaled by real body mass.
- 800+ cutouts across 400+ species, each cut from a real public domain plate. All art is historic and human-made. Coverage is currently best for the Nordics, Britain and Germany (but other parts of the world are in the works).
- Runs fully local on a Raspberry Pi or your homelab (yes, even the classifier runs great on a RPI)
The e-ink panel is optional and it can run web-only in a container against a BirdNET-Go you already have.
Happy to answer any questions!
Right now the coverage of North America is relatively sparse, but the interest has been high and I'm hoping for some contributions.
Edit: For North America the most obvious source is Audubon's Birds of America, which is public domain and should provide a lot of coverage. This collection includes some beautiful plates.
I'm based off Spain, so I'll be happy to contribute if I find missing Iberian and Mediterranean species, or if some cool feature idea comes up :)
Avian Visitors - https://news.ycombinator.com/item?id=48343424 - May 2026 (20 comments)
I can't help wonder what this would look like for humans. Imagine, set one of these up at a party and it has the means to identify anyone on the invite list. When it hears the person it displays a photo, brief bio, and relationship status.
The README.md mentions that "no art is AI-generated, though some has been retouched with AI." Superficially, that sounds mild. Some of the other closely-related projects, however, use different verbiage, or none at all.
For those of us with ethical concerns around generative AI application to artwork, much more detail may be needed before we could decide if this was appropriate for us.
Every bird is a cutout from a real scanned 1800s plate (Gould, the von Wright brothers, Dresser...). The manifest (https://github.com/arnegiacomo/fugleramme/blob/main/assets/a...) links to each source, so you can compare them to the original scans. No bird has been prompt/diffusion generated, although I did use diffusion-based tools to remove birds from the perches (https://github.com/arnegiacomo/fugleramme/tree/main/assets/a...).
Nothing at runtime uses AI either, at least in the sense of LLMs or diffusion. The collage is Pillow and numpy, and detection is BirdNET, which is a classifier, not a generative model.
However I have used LLMs for code-related work, and for writing scripts for programatically editing the images, like cutting, contrast and colour corrections.
Puzzlingly, no animal species requires AI, which is a mystery that need to be solved by AI.
should we tell them how hacker news gets onto their computer?