ESP32-C3 Adblock

(github.com)

146 points | by jayhoon 12 hours ago

20 comments

  • muti 11 hours ago
    Weird how the readme talks about hash collisions, but not in the way I would expect. Two blocked domains with the same hash isn't a problem, they both need to be blocked.

    Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com). There does appear to be a web dashboard and /unblock api so should be straightforward to resolve.

    • apefulsin 2 hours ago
      AI wrote the readme so it's unsurprising
    • zamadatix 10 hours ago
      Yeah, I think they have it backwards. As you say, regardless of how many entries you have locally, the collision risk comes from false positive hash matches not from worrying if the positive hashes collide.
      • 4gotunameagain 3 hours ago
        I think they have it right.

        If you have N entries where N >> 1, the probability of an arbitrary value colliding with the existing ones is Pcol(N+1) which is approx Pcol(N).

        • zamadatix 1 hour ago
          Edit: Skip to 4gotunameagain's reply, this probably doesn't contain anything helpful to read.

          The math is correct, but it's the math for the question:

            blockHashes.has(hash(domainInBlockList1) == hash(domainInBlockList2)
          
          While the thing they probably want to check the collision math for is:

            blockHashes.has(hash(domainNotInList))
          
          As that's the risk you do the wrong action due to a hash collision. Importantly, the 2nd calculation is not going to be bounded by the size of your block list, it's going to be bounded by the number of domains overall.

          The former question isn't really useful - you can have a 100 entry blocklist with 100 collisions and it still works the same as if it had 0 collision since the action on all 100 is the same. That's more a problem for generating unique IDs per hash whereas this is flipped around because it's about classification.

          • 4gotunameagain 1 hour ago
            No, it is for the right (second) question.

            If we assume uniform sampling for both the blocklist (size N) and the domain to be visited (not in the block list), the collision probability we care about is Pcol(N+1).

            The number of existing domains does not matter, only the size of the block list.

            Of course when we consider all 401 million domains there will be many more collisions, but in each case we only care about a collision in N+1

            • zamadatix 51 minutes ago
              Ah, I get what you're saying to do with the collision probability and that should definitely make sense but whatever they actually did in the readme ended up a factor of ~4 off the result that approach should give. E.g. the much simpler direct path of 2^40/537000 for that "next domain" question is giving me 1 in ~2.05 million rather than 1 in ~537k.

              Edit: Or maybe they formulated that path to derivation and how to map it back but couldn't quickly find a tool able to approximate the birthday problem to that scale so they tried direct testing instead? If, e.g., they were slowly upping the values to see when a collision occurred it would make sense they ran into a collision earlier than would be expected as they (effectively) gave more than 1 trial. Or just random chance too if that's really the path they took I suppose :).

              Edit2: After looking at the readme for far too long, I saw the original readme was in Japanese. Running that through a good translator actually clarifies or corrects a lot of the wording like saying "In the latter case, one other unlucky domain also gets blocked" rather than focusing on the number of domains over-blocked out of the 537000 only. With this version of the text I think you're definitely right, the math approach used should have given them the right number as the original text is already flipping things back to what happens with the one new query but they (apparently) just didn't do the actual calculation with it.

              Sharp catch :).

    • close04 4 hours ago
      > Two blocked domains with the same hash isn't a problem, they both need to be blocked.

      This is a problem if the second one is actually a major useful domain. If HN and adserver have the same hash, you have a problem. This could be described as a false positive for HN.

      > Where hash collisions matter is false positives, e.g. if hash(google.com) = hash(adserver.com)

      This sounds like the same issue I describe above where you need to use Google but block adserver. Otherwise it's only a problem if you implemented allow lists, right? This is when you explicitly allow Google and implicitly allow adserver along with it because of the hash collision.

  • yoavm 6 hours ago
    If you're ever thinking about getting an ESP32-C3, do yourself a favor and get the variant that you can connect an external antenna to. The normal C3 has a built-in antenna that is extremely weak, making it useless for most things I was planning using it for.
    • z2 1 hour ago
      And make sure to read the reviews as there are apparently fraudulent versions out there that don't include the 4MB flash chip. That said I took a gamble on a bunch of $2 C3 super minis, and they've had surprisingly adequate signal strength for indoor applications, even in the basement. (-60's dBm)
    • n8henrie 4 hours ago
      You can find instructions for making and soldering on a small antenna to the common "mini" dev board that reportedly helps quite a bit
    • utsavmishra25 2 hours ago
      yes and its such a important detail to have in later ESP projects. Having the ability to add antennas onto ESP-32 is a blessing-in-disguise that helps with more complex projects at a range later on
    • RicoElectrico 3 hours ago
      You mean chip antenna (one that looks like a big SMD resistor)? Yes, they're garbage. But this is orthogonal to the chip itself.

      Also, PCB antennas are reasonably good for most indoor situations.

  • BLKNSLVR 8 hours ago
    I'm a bit of a paranoid freak that likes lists, so I've got a PiHole that has a total list of 14-16M blocked domains.

    Great idea, and good for casual blocking, but I'm almost moving to an "allow list" mindset. This solution would probably work better for that, I wonder if the good parts of the internet would fit into an 140k list.

  • waysa 5 hours ago
    I think it could fit even more domains using a Bloom Filter or similar probabilistic data structure. With a chance of false-positives of course. But that's a trade-off the project already makes.
  • 1vuio0pswjnm7 10 hours ago
    Whitelist/allowlist is easier, e.g., it's smaller

    Depends on the user but not everyone is visiting new websites everyday

    Even for those that are, the number of domain-IP mappings needed will be relatively small

    Definitely under 140,000

    Most DNS data I use is "static", it rarely changes. As such most times I don't have to make DNS queries. I store the domain-IP mappings in proxy memory; this is faster than DNS

    No "blocklist" needed

    • calgoo 3 hours ago
      Just make a captive portal that shows up on your screen for new domains where you can just click "add to allow" or "allow this once" or "add to deny". That way you build the dataset, like most of these tools its annoying in the beginning while you generate the dataset, but after a while its only a few pages a week.
    • sheept 9 hours ago
      I would think that a regular user of Hacker News would be visiting new websites every day (though it'd definitely still be below 140k)
      • Etheryte 6 hours ago
        This is all a guesstimate, but my gut feel is that a considerable part of the HN population doesn't even read the linked content, only the comment section here.
        • jasonjmcghee 6 hours ago
          I sure hope that's not true. Maybe hit the comments section first?
          • dwedge 6 hours ago
            With github being the exception, I use comments to see if the article is AI. If it is, I prefer the condensed opinions in the comments. Not to say it can't be interesting I just don't want to waste time reading overly verbose generated text.
          • lhoff 6 hours ago
            Depends on the content, for all of these model release marketing sites, I for example only read the comments.

            If i would estimate it, I only take a look at 1/4 of the links where i read the comments.

  • timvdalen 6 hours ago
    > The trick everyone misses:

    Please just write the first sentence of your README yourself

    • dwedge 6 hours ago
      Apparently solving blocking extra domains due to hash collisions (reducing from 1 to 0) would be extra space "to solve a problem I don't have".

      I dislike when LLMs talk to me like that. I hate it when humans do it, confidently spewing their overconfident llm assumptions to others

      • lifeisloving 6 hours ago
        I actually thought this was one of the more communicative parts of the readme and quite liked their explanation.
    • WithinReason 3 hours ago
      The trick everyone misses: binary search of hashes can be done in log(log(n))
  • bilekas 3 hours ago
    Nice project and poc maybe, but I'm not seeing the practical application of this in a real world env, surely the latency is a deal breaker ?
  • anilakar 7 hours ago
    There's no point in using PlatformIO for ESP32 MCUs. The native ESP-IDF extension works much better.

    I would only recommend using it instead of the standard Arduino Processing IDE.

    • ricardobeat 2 hours ago
      I use arduino-cli exclusively, lighter than ESP-IDF and works completely standalone from the IDE.
  • jolux 7 hours ago
    10ms? Good grief that’s slow.
    • moebrowne 7 hours ago
      Still faster than using a remote resolver: https://www.dnsperf.com/
      • almog 2 hours ago
        But a remote resolver, such as the ones on the list, is not used to determine whether an item is contained in a set (0.0.0.0) but rather to retrieve the current IP of that domain.

        The way Pi-Hole is used is to first determine whether a DNS should be blocked and only if it's not in that set, forward it to a remote resolver. I'd think the local resolver should be at least an order of magnitude less than the remote to justify it.

  • fwip 11 hours ago
    Cool idea, latency might be too high, wish the docs weren't all AI vomit.
    • snailmailman 9 hours ago
      The biggest benefit of my local dns server is latency. On wired internet, my dns is <1ms from my PC.

      Upstream dns for me is pretty quick. Google and cloudflare dns are ~5ms from me. But WiFi latency alone is ~8ms most of the time in my experience. On my fiber internet, pinging a dns server in some random upstream server miles away is lower latency than WiFi 10 feet away. But the real issue on WiFi is any packet loss at all adding 50-100ms to that at random depending on interference.

      With DNS you are paying this latency cost all the time on nearly every request.

      • ktm5j 29 minutes ago
        But this is an ESP.. it's a 160 MHz microcontroller that does not have wired ethernet, only wifi. Also it's serving data from a USB flash drive instead of from RAM. Latency is going to be through the roof.
      • QuantumNomad_ 4 hours ago
        > With DNS you are paying this latency cost all the time on nearly every request.

        Surely not? macOS, Windows, and I think most of the big Linux distros, all cache DNS responses. Probably the web browser itself does too.

        • fwip 8 minutes ago
          Android and iOS do as well.
    • thenthenthen 10 hours ago
      This will be slow for one user, let alone more than one
    • danw1979 7 hours ago
      “The trick everyone misses”
  • rubyfan 2 hours ago
    I really like this conceptually. However, multiple elements feel like slop and that’s throwing warning signs to me. The Claude readme file and the tiktok-ish fast cut youtube video felt like Vince the Sham-wow guy put them together. Ironically it feels like the product of the problem we’re trying to avoid.
  • peter_d_sherman 2 hours ago
    >"The trick everyone misses: you don't need to keep the blocklist in RAM.

    Store the domains as sorted 40-bit hashes in flash and binary-search them.

    140,000+ domains fit in ~0.7 MB of flash..."

    Brilliant! (Although, what about potential hash collisions with domain names that shouldn't be blocked? Those are probably few and far between, so the utility of the block overall far outweighs any few false positives that might arise -- so I reiterate my claim of "brilliant"!)

    Also (if you're a compression nerd like I am), it may be possible to analyze this table of 40-bit hashes further to find the longest binary substring that repeats the most across all of them (or binary substrings that just repeat a whole lot), and design a Huffman-like (or other tree or tree-like) data structure around those, and possibly compress the data further...

    Of course, then the simplicity and elegance of the Binary Search might be lost, but it might be an interesting optimization exercise to know just how far that data could be compressed... "Hey Claude, try to optimize the compression of this data using alternative compression data structures, and report back to me!" (Or something like that! You know, get the AI coding agents to try different things -- or attempt to code it on your own (always better for understanding!), etc., etc.!)

    Anyway, brilliant!

  • jwillmer 4 hours ago
    pinhole is great but if the device fails your network is down. I swapped to nextdns because of that. maybe support of multiple devices would fix it - now that is this cheap it would not matter .
    • fwip 4 minutes ago
      I have my router set up to send DNS requests to my adguard home server, and if there's no response, it falls back to one of Cloudflare's. So it fails-open, letting the ads in along with normal requests.
    • import 2 hours ago
      You can have a 2 adguard home + adguard-sync. Works smoothly.
      • b3lvedere 2 hours ago
        But what if both fail? :)
    • close04 4 hours ago
      The simple answer is have 2. If the cost per unit is low enough than deploying 2 PiHoles is trivial.

      The developers could make users' lives easier by implementing a clustering option that syncs the 2 out of the box. For static deployments, where you don't change the config too often, it's still decent as it is.

      • krtkush 4 hours ago
        I did the same (until my RPi zero stopped working). Both the devices share same vertical IP address and if one ever went down, the other would take over without any delay.
  • nicman23 8 hours ago
    i dont get pihole. just use a proper dns? if you want local, just use unbound?
    • Tajnymag 7 hours ago
      Pihole and unbound can work together. Pihole isn't a good DNS server, it's a good adblocking DNS server.
    • dannyw 7 hours ago
      people like ease of use, UI, docs, and community.

      after all, why use opnSense; just use suricata and $PROPER_X! why use TrueNAS; just use $DISTRO and ZFS!

      • nicman23 4 hours ago
        it is easier to add adguard dns if you want ease of use
  • rekoil 7 hours ago
    "They were too preoccupied with whether they could, they never stopped to think whether they should"

    Jokes aside, very impressive that it works!

  • manlymuppet 8 hours ago
    Wow, this is an atrocious name for a project haha. Cool though.
  • shevy-java 3 hours ago
    I have a weird opinion: I believe adblocking must become a basic human right. Conversely Google trying to deny ublock origin and other extensions, should be fined many billions of euros as punishment. Then be disbanded if unable to pay. That money also should directly go back to The People, since they suffered from this evil malpractice of Google and others. You can see that currently the opposite is happening e. g. AI slop companies pushing ads down to people. That's why we must make adblocking a basic human right - free access to information at all times.

    It may sound strange right now, but people in the future will ask why we did not push back against these evil leeching companies before. Those fines are no longer enough nor sufficient - CEOs that break human rights need to go to jail for mandatory 5 years minimum, without any way to weasel out their way via money. Right now money is a way to get out - that system is broken. Anyone wondered how Epstein got so much money in the first place? Hint: he did not build it up by himself only.

    • ricardobeat 2 hours ago
      Let’s not trivialize “human rights”, those are much more fundamental issues and this kind of discourse does not help achieve anything.
    • cindyllm 3 hours ago
      [dead]
  • nazmul_ai 3 hours ago
    [flagged]
  • aiXis 7 hours ago
    [flagged]