Except that the reason that the YouTube “1A Auditors” get the attention that they do is that frequently it’s not a simple “check things out, and inform the caller that the person recording has the right to do so”, but the interactions often escalate to threats of unlawful arrest as police default to the side of the caller.
Calling the police, especially in America, can 100% be a coercive use of force. That’s why people threaten to call the cops/DHS/ICE all the time during arguments, and why things like SWATing are as big of a problem as they are. When the behavior of many policy officers is escalate first, ask questions second if at all, it does become a genuine threat.
I don't know what to tell you. I watch these videos all the time (it's one of my vices, bodycam footage in particular) and what I see every time --- when it's "auditors" harassing businesses, not when it's auditors confronting police at police departments --- is business owners being told to suck it up.
I know that in Summer 2020 there was a fashion for saying any call to the police was ipso facto a deployment of coercive force, but I'm rejecting that argument out of hand while also pointing out that in many regions 911 is exactly the thing you're supposed to do when you have uncertainty. It is simply not a demand for violence; it's a demand to have a trained authority come and resolve a situation, rather than having two private citizens with no training at all in the law confront each other and hope for the best.
GitHub Copilot engineer here working on identity, safety, and privacy - no, even Microsoft doesn’t have access to all GitHub repos.
As years have passed since the acquisition “company” delineations have blurred a bit, but Microsoft employees still need to go through a separate onboarding process to access any GitHub company resources (internal repositories, telemetry, documentation, etc.), and then we have an additional layer of entitlements to gate and audit access to any sensitive data, including user data.
Very few employees within GitHub proper even have access to view private repositories, and in the rare cases where that’s done for legal or safety reasons the repository owner is notified.
There are currently no OpenAI employees with access to GitHub systems, so there’s about 4 layers of protection in place to prevent private repositories access. We do genuinely take user data protection and privacy seriously.
This is a nice answer to the question "how is GitHub preventing rogue employees at Microsoft from stealing my private repositories?". Like, it's good to know I'm covered if Microsoft accidentally hires a North Korean spy or something.
But if Microsoft really was selling private repo content to OpenAI, it probably wouldn't go through those access controls. It'd be an executive-level decision with enough force to plow through all the red tape, and it'd be implemented as a data pipeline or similar automated process that wouldn't trigger the same kind of notification as, like, a Trust and Safety employee taking manual action.
Probably the better evidence here is in GitHub's ToS where they say in pretty strong/binding terms that they aren't doing this: https://docs.github.com/en/site-policy/github-terms/github-t... . If they are secretly selling your data to OpenAI they haven't left themselves a ton of wiggle room if people ever found out.
(Probably the biggest loophole they could use is to send private repo content to an OpenAI service for scanning/safety purposes. The ToS allows this and they're almost certainly doing it with other services like PhotoDNA. Then OpenAI can just violate whatever agreement they have not to store the data sent to that service.)
I’m one of the people directly responsible for ensuring that those terms are properly enforced. Presently I’m arguably the person for Copilot data specifically.
Current talk of the town in the data retention space is around AI safety. There’s been a recent slew of blog posts and academic papers around how LLM harms can manifest over multiple agentic turns, from individually innocuous requests. Identifying this inherently necessitates user data retention which we do everything possible to avoid (not even meaning data sharing as is alluded to in this thread, I mean literally persisting prompts and completions anywhere outside of ephemeral memory). I’ve been the one advocating for having the storage of any data retained for safety and security purposes to be as heavily access controlled and audited as is possible.
Also, if AI safety is a space that is interesting to you, we’re hiring! Manager, developer, and applied science roles, or we can figure out the HR shenanigans if you don’t fit any of those archetypes. If interested shoot me an email at taywrobel@github.com!
(FWIW, in my conspiracy theory, the data sharing would be buried in the part of the company responsible for making sure that people don't upload e.g. CSAM to private repositories, so the Copilot people wouldn't be directly aware of it. I might've edited that in after you already started writing your reply though.)
Still in my bubble! I am not involved in the human review or automated analysis portions of the safety pipeline for CSAM/TVEC harms, but my team is responsible for the data handling around identifying and responding to such content.
As of 11 days ago our vision support is GA (https://github.blog/changelog/2026-07-01-copilot-vision-is-g...) and let’s just say the technical implementation wasn’t the long pull there. Figuring out the what and how of responsible data handling around what I hope is agreeably harmful use was… quite a journey.
> Very few employees within GitHub proper even have access to view private repositories
so we're just discussing what business Microsoft likes more at any moment. and you didn't provide a list of allowed use cases (is Ai training one?). making your huge answer(s) empty and not contributing one yota. sorry.
i feel your job exist to uphold the illusion and you will not see it any other way.
How do you define "access" here? Microsoft has demonstrated that it can delete any GitHub repo at will. Maybe there's some shell entity between corporate "Microsoft" and "GitHub" that's doing the dirty deeds without attribution...
Access meaning read, modify, delete, etc. Pretty standard definition, unless you know of a different meaning of access I’m not privy to.
Microsoft can certainly request that we perform actions against repositories, as can governments, customers, random people on the street, etc. Whether action is taken in those cases is a question for lawyers to fight over, but we have the engineering guardrails in place to require it to be an intentional, audited action.
I appreciate the spicy question tho, even if misguided!
Even with your rephrasing you’re looking for an answer in absolutes which is generally impossible, but unpacking your line of questioning, what it really amounts to is how “in the know” I am or am not.
To the best of my knowledge I know about every ongoing company AI safety and user privacy initiative, and none of them involve permitting access to copilot user content to any second party or third party entity.
Of course, that’s tautological. I don’t know what I don’t know, but I’m senior enough and with broad enough scope that I’m at least read in on what I believe is the majority of high level business initiatives.
I’m not trying to be evasive, this is just the reality of any organization - I only know what I know. Everything within my scope of awareness indicates that there is no copilot user content access outside of our publicly published terms of service.
Thanks for responding. It's great to hear from someone working these issues day to day, and it's the reason I come to HN. I feel like this particular line of questioning is a bit silly, with all the "Can you absolutely guarantee X, Y, and Z?" Thanks for engaging despite the adversarial turn it has taken!
I think there's maybe a disconnect here that people are largely concerned with the contents of their private repos, while you're maybe more familiar with how AI interaction data is handled. (After all, the original topic of the thread was X.ai allegedly going above and beyond interaction data to exfiltrate entire repos.)
I personally did get the vibe that you were being evasive, just because the things you were saying didn't quite match what people were asking about, in a way that felt kind of like a corporate legally-not-a-denial denial. It's like, "Hey, has Contoso Apartments hidden a camera in my bathroom?" "Contoso Apartments is committed to your privacy and safety. We have strict controls in place to ensure that our maintenance staff cannot make a copy of your key without notifying you. To the best of my knowledge, we do not have any company initiative that involves opening envelopes addressed to you." Like it's theoretically reassuring for the company to commit to those things, but the fact that they can't directly answer the original question is disconcerting.
Ultimately there's probably not a whole lot you can do about this. Like realistically if Microsoft is doing this, they've probably constructed it in a way where not many people know and/or they can plausibly deny it. So it comes down to (a) Microsoft denies doing it, but isn't making the broadest legally binding commitment possible, (b) does the reader believe Microsoft and OpenAI are trustworthy with respect to privacy and intellectual property issues or not.
I've been in this kind of situation before, and it can be frustrating when people don't believe that you're in a good, isolated department of the company and you're committed to upholding ethical standards. I guess that's why big companies pay the big bucks :)
That’s fair, and I appreciate the more constructively critical feedback! Also worth noting that I’m solidly on the platform service side, and the article here is largely focused on malicious client behavior.
Copilot has a lot of different clients between IDEs, agentic integrations, and GitHub apps. I don’t have awareness of the implementation details of all of them, but I can assure you that we don’t provide APIs like those mentioned in the article being used for data exfiltration.
Clients are responsible for context building, and all go through the same service that does auth, policy and quota enforcement, request routing to the underlying providers all of which have zero data retention enabled unless very specifically excluded from that (looking at you, Fable 5).
> If we lower the threshold from "absolutely" to "absent third-party breaches" what would you say?
If anyone answered that question affirmatively, I'd lose a massive amount of trust in them. It would betray they fundamentally misunderstand the stochastic nature of playing defense.
I do not believe in being non-antagonistic in pursuit of truth. They could have admitted they never had any control over what's uploaded. Instead even when "antagonistically asked a question" they chose to not answer it. This is an answer in itself and to our great sadness extracting that required some amount of "antagonism".
Prove that I work at GitHub? Username + LinkedIn can show (not prove) that easily.
Prove that we have an entitlements system which regulates and audits access? I could point you to https://github.com/entitlements, but it’s all private repositories so that won’t prove much either.
Prove that there are no OpenAI employees with access to GitHub systems? Not sure how I’d do that without dumping (what you would still need to trust me is) the entirety of our org chart/HR system, which I’m not willing to do because I do enjoy being employed and am not exactly obfuscating my identity here.
Prove that HN has a strong anti-Microsoft bias? Well that one is pretty easy actually, you’re helping prove it yourself!
Let’s be real, we now live in a post-truth world. Nothing can truly be proven or disproven outside of formal logic and mathematics. You can either believe what I’m saying as good faith insider knowledge sharing (which is unfortunately rare nowadays) or you can not. Makes no difference to me.
That is my point! The fact that there are a lot of people that will believe what you are stating without a reasonably proof is what makes me sad and worry about the future. That kind of statements you are doing are enough to put in company webpage and term of service and thats it. Any attempt to repeat them as if they are true makes:
1) The messenger looks good in the eyes of stupid and innocent people.
2) The messenger looks stupid in the eyes of people that have reasonable doubts about company statements that are agains their own interest.
Without robust and easily scaled infrastructure in place ahead of time, an organic DDOS is one of the most difficult situations to mitigate. Not much can be done in terms of traffic shaping, rate limiting, or bot detection.
An HN front page “DDoS” is like 20K hits. This isn't some complex scaling challenge. Any website on the internet should be able to handle it, especially a purely informational one.
I had my blog be on the front page for ~6-8 hours racking up 100k+ unique loads. It also managed to survive just fine on a $5 VPS so I would hope that other sites could survive.
I agree. Protecting against DDoS attacks is incredibly difficult. I'm just enjoying the irony of Def Con, the premiere computer security and hacking convention, not being able to handle traffic.
To be fair, I don't think they crashed; I saw a "sorry too much traffic try later" type message. Still amuses me.
I guess it's funny, but the attendees don't necessarily represent the organizers. The best hackers in the world may be in the building during Defcon but I don't think the Defcon organization itself necessarily employs them.
the current way to most effectively get around DDoS seems to be using a proof-of-work based frontend run on as many revolving reverse proxies around the world as you can afford. this is what kiwifarms does. seems pretty effective and a lot cheaper than what the people bankrolling the attacks on them are spending.
Wow, I was at Apple back in the 2018 timeframe when Peter was first building this. He was hoping to make it open sourced even back then, 6ish years ago. Great to see that it finally made it.
I really wish Apple would learn to play nicer with the OSS community. I have yet to see them deciding to open-source something backfire on them monetarily or reputationally, and I've seen the act of them abruptly close-sourcing things sour community opinion (i.e. FoundationDB).
There’s a reason that this is called “hacker news” and not “just use the industry standard for the last 3 decades news”.
Won’t downvote you for giving pragmatic advice, but I appreciate projects like this that slap together disparate technologies for an interesting goal, even if it isn’t the best choice for your usual Fortune 500 company.
If anyone else is as frustrated as I was with the article mentioning “the DAW” 73 times without defining once what the actual acronym stands for, it’s “Digital Audio Workstation”.
In the same way I don’t expect a biologist writing for biologists to explain “DNA” stands for “deoxyribonucleic acid”, it’s probably not necessary for a music producer writing for producers and engineers to define “DAW”.
Users here probably feel the same way about HTML, FIFO, DAG, etc
Yeah the top 100 is super weird, it’s all these commercial EDM DJs but the weird thing is the magazine doesn’t otherwise really seem to target that audience. I don’t read it but I have come across some good long form pieces like this from them online, so actually I think they are trying to do some good stuff.
I'd imagine that it's because the DJs ask people to vote for them. A lot of the DJs I follow do that every year. If the big commercial DJs with the biggest following do that, then they would naturally land at the top.
Yeah I guess what I mean is that it seems to go against the rest of their brand, the magazine usually covers slightly more underground dance music it seems - not super underground, still big names, but not stadium EDM stuff.
Maybe their philosophy is that the Top 100 should be an open thing and they shouldn't restrict who can enter based on music style... to me, it makes DJ Mag way less credible, but I guess they probably make money out of the Top 100 being so big.
> +1 Feels like they don't care who's their readership. Felt like they told me: "If you're not in the industry, Google it."
Caring about their readership is exactly what they're doing, just that you happen to not be what they think of when they imagine the typical reader. The typical reader is already into music production and with a 99% certainty know what a DAW is.
I wouldn't expect every tutorial on "Google's Official Android Developer Blog" to explain that "JVM" means Java Virtual Machine, some resources really are for people who already know a bit about the subject area.
I wasn't in this instance, but am in general. For industry folks they probably don't even realize it's not a word--surprised it hasn't lowercased to daw by now /s.
On the one hand, I've seen people (including myself) try to hack job-queue like semantics onto Kafka many a time, and it always hits issues once redelivery or backoff comes up. So it's nice to see them considering making this a first-class citizen of Kafka.
On the other hand, Kafka isn't the only player in the queue game nowadays. If you need message queue and job queue semantics combined (which you likely do), just use Pulsar.
I think the most likely use case, the one making me happy they're working on this, is reducing infra spend and having a separate tool/guarantees/storage for queues and for whatever kafka is more made for.
I'm just hoping librdkafka gets good too-tier support for this feature in a timely manner.
Calling the police, especially in America, can 100% be a coercive use of force. That’s why people threaten to call the cops/DHS/ICE all the time during arguments, and why things like SWATing are as big of a problem as they are. When the behavior of many policy officers is escalate first, ask questions second if at all, it does become a genuine threat.
reply