Make no mistake - as far as I can tell, this preserves Google's ability to track that you are, say, a new parent because you've searched for baby clothes. What it won't know is which new parent you are. The system is designed to give probabilistic assurances of k-anonymity. But Google will no doubt tune those "cohort" memberships, and the value of "k," to capture the vast majority of current advertiser needs, while still being able to communicate to antitrust inquiries and the public that they are not giving people unique identifiers. If anything, it hurts their competition more than it would hurt them, because it allows them to thread the needle in a privacy-conscious world.
Still used to the 90s when any Level of bypassing would get you labeled as a hacker, and people would demand you get fired or arrested.
Just knowing how things worked resulted in a number of people assuming I could hack banks, and either asking me too, or shunning me. Or flat out accusing me of it.
Umm you could go around mugging people much easier then I could hack stuff. That explanation never went over well.
Had to ask one boss to stop threatening employees to have me hack them. It was in jest, but I was terrified at how people would react if they thought I bypass security. Didn’t help that I knew how, as security was crap back then.
I leaned to keep my mouth shut, and refused to keep up with security stuff. Can’t hack if I don’t know how.
To each her own, but I see the complex browser such as Chrome, Edge and Mozilla as part of the problem. Javascript too. Neither is inherently problematic, but both are too often used to create problems for users (to serve advertisers' interests) rather than to solve them. For example, without Chrome/Edge/Mozilla plus Javascript, "paywall" does not work. Yet I can still read the WSJ article just fine with either Chrome/Edge/Mozilla nor Javascript.
Much like a tech company that wants www readers to believe it is working to solve a problem www users are having that the tech company itself created (which, it just so happens, is not a problem at all for advertisers).
> Make no mistake - as far as I can tell, this preserves Google's ability to track that you are, say, a new parent because you've searched for baby clothes. What it won't know is which new parent you are.
If you've searched for baby clothes in Google there's absolutely nothing in the FLoC proposal that stops Google from inferring that you're a parent from that search and, idk, storing a bit in your profile or something.
This proposal is about allowing ad targeting to interests or whatever without tracking your visits across the web - your browser generates the cohorts clientside (idk how the cohort assignment algorithms get standardised on, I guess that's where the 'tuning' would go). And yes, you could just tell your browser not to do that, or have it only use the most recent n days' data, or randomize your cohort, or freeze your cohort in time, or blocklist certain sites from ever entering your cohort, and so on.
Where this ties into tracking if you're a parent is that while it doesn't prevent tracking that, it does lessen the incentive - if you're a parent, you're likely going to wind up in a parent-y cohort. Is the signal from the original search still going to be useful enough to warrant keeping around?
If the browser is supposed to send the information, what are they going to do with those of US who do not use Chrome? Are there plans to force firefox to follow the plan, and how will they prevent others from, say, making an extension to mess with the cohort numbers?
I think the expectation is other browsers will implement by default - the intention incidentally is that FLoC isn’t a google specific tech, it enables private targeting by any ad vendor. The proposal seems to mostly to be by googlers but the association is with WICG which seems to have people from all over.
As I noted in my original message - yes, its client side, you can do whatever. (‘Whether the browser sends a real FLoC or a random one is user controllable’ https://github.com/WICG/floc)
The user controllability is genius, that neatly ties up their problem. They have a legitimate option to get out of being targeted, which they know almost no one will ever use. They could probably even add a setup prompt on install that says something like "Let us know which of these categories apply to you, so we can tailor your browsing to your interests" and get people to fill out half of it for them so they don't have to guess.
> "Let us know which of these categories apply to you, so we can tailor your browsing to your interests"
This prompt will never happen. The last thing Google wants is any additional association between their brand and advertising.
Google's mode of operation has always been to sneak in the back door and count how many boxes of cereal you have while you are looking at underwear. Knocking on the front door and to ask if they can come in would likely never occur to them.
There is no mention of Chrome being an ad supported product anywhere when you install it. They want that association completely out of people's heads.
This is a misunderstanding of what these cohorts are. If you read the [FLEDGE readme](https://github.com/WICG/turtledove/blob/master/FLEDGE.md) you can see that there are multiple types of owners of cohorts that fall into 3 categories:
1. Advertisers - who would add you to a cohort they own, so say "Nike - womens-running-shoes"
2. Publishers - who'd want these cohorts to better allow for advertisers to advertiser on their site, so "NYTimes - business-section-reader"
3. Third party ad-tech companies looking to create audiences for advertisers who work with them, so "Example Agency - mens-formal-wear". They'd partner with publishers so that when you go to somewhere like GQ or Man of Many and read an article on tuxedos then you'd get added to the cohort.
So you are right, you'd never see a prompt to add in your interests, but this isn't because Google doesn't want to associate with specific brands, it's that Google isn't necessarily an owner of a cohort (though they totally could create their own under any 3 of those categories).
Cookie popups are now part of our superstitions, our Internet rites. A decent web page, it doesn't just track you and molest your integrity -- it asks first!
I don't think we'll get away from the cookie popups, ever. It's just like the "I agree to the TOS / EULA" checkboxes. Legally toothless, but you'd better have one, just in case.
And cookie popups are now part of any corporate general counsel's requirements if you're even running first-party cookies and/or any kind of analytics, regardless of third-party-cookie changes. They're not going anywhere.
I disagree. There is are many knowledgeable people out there that spend their time looking at the requirements of GDPR etc. and how those relate to cookies. If you can be properly advised that the kind of cookies you use aren't contravening the various privacy directives, it won't be long until they disapear.
I know one place where they were forced to add a cookie popup as part of a GDPR compliance tickbox exercise. The kicker? They didn't have any cookies. After a fruitless attempt to persuade the person working through the aforementioned exercise that you only need a cookie popup if you have cookies, the popup was implemented.
Of course you need to store the user's preference. So now the site needed to have a cookie. One cookie which just stores the fact that the user has or hasn't consented to cookies.
It could have but because DNT wasn’t made part of the GDPR and sites refused to do the right thing on their own it went no where.
Even cookie pop ups are not respected. If you visit Facebook sites by accident they’ll leave their tracking cookies before you’ve had a chance to read the terms and leave the site.
I know. But it doesn't matter as DNT is universal. If you enable DNT, websites SHOULD (or better: MUST, but unrealistisc I know) honor that. For technical/functional cookies you don't need permission. The rest is off, DNT answers the question of "Do you accept" with a browser-based "No.".
Chrome isn’t critically important to Google’s advertising and data initiatives. It’s important for a lot of other things, but that’s beside the point.
Google doesn’t really need much more that what it has. It has third-party big data, Gmail, Google Search and a very recent optimized cache of everything important that it’s crawlers can get and more. People get so caught up in the minutia; Google is big, like category 100 Hurricane big, where people laugh and say there is no category 100; but that’s what it is- classification system aside.
Anti-trust and pro-privacy are distractions that Alphabet, Google, etc. can put their left-fielder on while they run the bases, because they play both teams.
Presumably people would just block the ads instead?
Either way I reckon it's the same thing stopping people from just sending random data over google-analytics. So, not a whole lot, other than whatever anti-spam mechanisms google bothered to put in place.
>> just sending random data over google-analytics.
Back when I was in the security field, I heard discussions about this sort of thing as a form of attack. The question was whether a computer that sent false or incorrect information was either engaging in a DOS attack or was violating the CFAA in that it was using false information to access a service. The DOS attack would be on the tracking system rather than a website resource, false information to degrade the effectiveness of the tracking system overall. The CFAA angle would be that participation in ads/tacking was part of the contract for accessing the 'free' service, that by blocking ads or sending false info the user is using a false credential to access the service. At the time, many thought that by using an ad-blocker that users were failing to pay, passively stealing website content. Those who sent false information were engaged in active theft.
I'll grant that sending fake data is more active than merely not loading ads, but I personally consider not loading ads about as malicious as failing to read a billboard, or failing to subscribe when a webpage says 'Subscribe now'.
I don't think I'll ever quite understand the idea that the user is under any obligation to do whatever. Just because their user-agent might have some default behaviour doesn't mean that the user-agent should be expected to act in the interest of the server.
Though there's an argument to be made that sending lots of false data in response to a request to post personal information to google-analytics is a bit like manipulating a public survey by sending in silly answers. I'm not sure to what extent that is illegal but I'll grant that it's not exactly ethically correct.
>> not loading ads about as malicious as failing to read a billboard
Because, at this time, ad blocking tools are relatively benign. But what if we think of them more as content-blocking or content-management tools? What if I have a tool that blocks only right-wing advertisements or content? What if I have a tool that blocks all images of women showing too much skin? Block images of black people? Such not-neutral tools would radically change public perceptions. Legislation might follow.
There was a comical toll out there that would replace the text "Trump" with "The idiot" on viewed websites. That's the thin edge of this wedge.
This is approximately the time when you go from passively not agreeing to be tracked or have your attention stolen or your computer compromised by ad networks to collaborating with your fellow citizens to destroy the business of the ad networks by making it difficult to impossible to operate. In truth if every free ad supported service ceased to exist the internet would go on none the less.
I'm deeply interested in these issues as well - the criminal aspect of the CFAA being disconnected from actual harm caused to people is insane to me.
There does seem to be momentum towards a narrow reading of the CFAA when it comes to providing false information against terms of service, though, as per these 2013 analyses... though they fall short of actually giving any guarantees.
On the other hand, a denial of a motion to dismiss in Ticketmaster v. Prestige from 2018 seems to have drawn a distinction between breaching terms of service alone, and being told in a cease-and-desist not to breach those terms; the latter more clearly outlines what would be considered exceeding authorized levels of access.
Where does false information end and privilege escalation begin? Can that question depend on whether your privilege escalation is something as simple as "I'm now able to avoid the fine-print contract that I need to watch certain types and customizations of ads in order to access the service?" What constitutes sufficient notice to a user that certain actions are explicitly forbidden, if a C&D does but terms of service do not? All questions that, as far as I can tell, haven't been fully answered.
(Obligatory: Not a lawyer, the above is not legal advice.)
> The CFAA angle would be that participation in ads/tacking was part of the contract for accessing the 'free' service, that by blocking ads or sending false info the user is using a false credential to access the service.
I'm horrified whenever I see people having this perception of the issue, and worried that it'll get accepted by the legal system as the correct view.
From my POV, the only contract that exists between me and a random ad-loaded website is the one negotiated between my computer and their server - that is, the HTTP protocol, which clearly stipulates that I can render any reply you send me in whatever way I like. This contract provides many tools for expressing the intent of gating content behind some requirements (like payment, or receiving different content first), but the common rule is: you don't send the content you're gating before the client proves they've met the requirements.
And honestly, I'd consider mixing ads into content to be borderline CFAA (the ads themselves being abusive and often fraudlent, and sometimes malware) - and sending software intending to compromise my user agent (like trackers, or ad-blocking busters) to be crossing the CFAA line. Unfortunately, I don't think the lawyers would agree with me :(.
Are you horrified when a movie theatre doesn't want people sneaking in their own food and drink? Some consider that trespass in that you have broken the contract you agreed to when buying the ticket. They can call the cops and have you arrested for accessing the theatre in contravention of the rule. I could see an equivalent for websites that adopt a "no ad blocking" rule.
If you are asked to leave and do not leave I can see you possibly being arrested for trespassing, but I fail to see what you would be arrested for by bringing in food from outside the theatre. Sure, they could ban you from coming back to property (and this could lead to trespassing arrests, if you did return) but thats about it.
I'd expect such websites to communicate it, though, and do the minimum of work to prevent me from entering with ad blocker on. I'm fine with paywalls, and even "disable your ad blocker" walls - they're up-front, and in agreement with spirit of the HTTP protocol.
Much like cinemas. If I try to bring in food I didn't bought at the cinema store, I will be stopped at ticket check and refused entry. Also, the cinema doesn't require me to buy food as a condition for watching the movie. If they did, I'd consider it the price of entry.
What current ad-loaded websites are doing is the equivalent of letting you in to watch a movie, and then having someone come to you as you're watching, and poke you and yell at you because you didn't buy anything at the store.
> They can call the cops and have you arrested for accessing the theatre in contravention of the rule.
I've never heard of a cinema calling cops on people who snuck in food into the theatre. Is that US-specific?
It only happens when people are discovered, begin arguing, are told to leave, and don't. It's mostly drunk people who have brought alcohol into a movie. Cops remove them as trespassers who refused to leave after being told to.
I might imagine that Google Inc could sell a product to Mozilla Inc which they could embed in their Firefox browser, enabling them to sell FLoC data to Google. For example.
> If the browser is supposed to send the information, what are they going to do with those of US who do not use Chrome?
What would stop them from implementing FLoC in the cloud for non-Chrome users? Fundamentally, where the data is stored is meaningless. So long as they are pushing everything through this same "FLoC" model, they are claiming they won't be tracking individuals.
So on Chrome, FLoC is local—your own browser working against you—for Safari and FireFox, FLoC is cloud based.
So long as the web sites in question are running Google Analytics, they still have the software on your system on a huge number of sites.
This would require being to identify a user when they are on many different sites (cross-site tracking), which all the major browsers intend to prevent.
(Disclosure: I work at Google, speaking only for myself)
First, you have to understand that Google as a company has lost all credibility in much of the community. There have been too many issues like this one where Google gives you the impression they are turning off tracking... but doesn’t actually stop tracking you.
There was also the Do not Track setting, which Google ignored. So in my book (and I’m not alone here), anything coming from Google is taken with a huge grain of salt.
For some time, there has been a sort of cat and mouse game between advertisers and browser makers. Do you know Google isn’t using fingerprinting to identify us? Or just IP or some other means?
> Google as a company has lost all credibility in much of the community
What I'm saying is Safari [1], Firefox [2], and Chrome [3] are planning technical mitigations that would prevent anyone from implementing a cloud-based FLoC. Even if you don't believe Chrome, it wouldn't be possible in Safari or Firefox.
> Do you know Google isn’t using fingerprinting to identify us? Or just IP or some other means?
https://blog.google/products/ads-commerce/a-more-privacy-fir... has "once third-party cookies are phased out, we will not build alternate identifiers to track individuals as they browse across the web, nor will we use them in our products" and "our web products will be powered by privacy-preserving APIs which prevent individual tracking while still delivering results for advertisers and publishers" which seem pretty clear to me? Additionally, I work in client-side ads infra, and I'm pretty sure I would know if Google were fingerprinting users to target ads.
From what I've read on FLoC (not much so I may be wrong), wouldn't having anonymized ad-targeting be a good outcome? Google and websites still get their ad revenue, and we still get ad-supported products. I guess there's always the possibility of a badly-anonymized algorithm leaking data, but other than that it still seems like a step up from that invasion of privacy being a given, no? If I'm misinterpreting the jist of FLoC, let me know.
I think an important question with the FLOC stuff is sort of brushed aside.
So, you have millions of users, and you make it "k-anonymous", or "anonymous if you squint real good". Is it really possible to find k such that privacy is meaningfully preserved, given the heaps of other information you have?
Say I belong to some cohort, and you know my IP address. I think this would be enough to fingerprint me reliably, never mind OS version, browser, plugins, etc.
Even if they stay with cohorts, I assume people will be members of a variety of them.
New parent. Lives in Springfield. Works at a power plant. Drives a car. Owns house. Horrible credit rating. Married with 2+ kids. Voted in last election. Ex military. Ex astronaut. Once purchased an effective baldness cure. How many seemingly large cohorts does it take before you are really talking about one clearly identifiable person?
> How many large cohorts does it take before identifying a person
It takes 33 “perfect” yes/no questions to identify uniquely anyone on earth. It is seemingly large since each question splits the world in two.
Of course this is in a perfect world, but I remember seeing quizzes which would guess any object given a few yes/no question. Or perhaps it was people.
> Of course this is in a perfect world, but I remember seeing quizzes which would guess any object given a few yes/no question. Or perhaps it was people.
Couldn't get to Dianne Morgan in 40 guesses. The closest it got was to ask in sequence whether the character was British, blonde, an actor and a comedian.
Not many probably know the person in question with her real name - but almost everyone who's ever watched Charlie Brooker's productions should recognise her as Philomena Cunk.
"Each player starts the game with a board that includes cartoon image of 24 people [...] Players alternate asking various yes or no questions to eliminate candidates."
Err, what? So there would be one cohort of "new-parent-lives-in-springfield-works-in-power-plant-etc-etc" (combining multiple unrelated interests), it just has to have thousands of people in it?
Do cohorts allow advertisers to say, "Show this ad to anyone in a cohort with an affinity for 'new-parent' over threshold X?" is that so?
Do they let advertisers say, "here are the cohorts of our best 1000 customers, please show ads to anyone else in those cohorts?"
Would people in "my cohort" be the 1000-10000 people most like me in browsing habits in the world?
(Disclaimer, have tried to buy limited quantities of ads in the past, interested in some part as an advertiser)
> Would people in "my cohort" be the 1000-10000 people most like me in browsing habits in the world?
Yes: "a browser can group together people with similar browsing habits, so that ad tech companies can observe the habits of large groups instead of the activity of individuals. Ad targeting could then be partly based on what group the person falls into."
> there would be one cohort of "new-parent-lives-in-springfield-works-in-power-plant-etc-etc" (combining multiple unrelated interests), it just has to have thousands of people in it?
I mean, there aren't thousands of new parents who live in Springfield and work in a power plant? You would be in a cohort like "43A7".
> Do cohorts allow advertisers to say, "Show this ad to anyone in a cohort with an affinity for 'new-parent' over threshold X?" is that so?
Not really? An advertiser or (ad tech company) might learn, over time, that people in cohort "43A7" are very likely to be interested in baby clothes while people in cohort "5B7E" are more likely than average but not by that much. So they might be willing to bid more to show baby clothes to someone in 43A7 than in 5B7E, and not bid at all for most other cohorts. But the browser API just gives you an opaque cohort identifier.
> Do they let advertisers say, "here are the cohorts of our best 1000 customers, please show ads to anyone else in those cohorts?"
From my reading of the spec, that seems like it would work well.
Could a browser vendor send down a mapping of cohort hashes to interest vectors? I don't see anything in https://github.com/WICG/floc that suggests that couldn't happen, and it seems like it could greatly enhance utility without too much sacrifice to privacy. For example, it'd allow an advertiser to target all cohorts who like american football, I'd think.
I think this can be built on top of the proposed API? For example, a site about American football could publish/sell the cohort frequencies they observe.
While I completely agree with your point, and the number required is probably pretty low, the point of FLOC is that no one will need to identify a real person for any “proper” reason.
Advertising rarely is about targeting “John doe” but usually about targeting “[men] && [in Springfield] && [searching for a baldness cures]”.
(I still will do everything in my technical power to disable floc, and still will never trust google). But I do think the implementation removes most of the advertising incentive to track on an individual person level.
Legitimate advertising. What about the other stuff? What about advertisers who want to swing very specific voters? What about nefarious people who want to put a malware link in front of a specific person. [active military]&&[over 55]&&[lives beside base X]&&[awake before 6am]&&[college education]&&[searched for "retirement planning"] will probably get you the most senior officer at a base/unit. Same too with senior politicians.
> Legitimate advertising.
> What about the other stuff?
Thats what makes FLOC tolerable. The legit stuff is enabled, and it makes it harder to get the other stuff (not impossible of course).
> What about nefarious people who want to put a malware link in front of a specific person.
It'll at least be harder than pre-floc (ideally).
The point is to remove the incentives to collect as much data about an individual (you don't need full profiles of someone) by creating even easier ways to get useful targeting without the direct individual identification.
---
(also, i'm broadly against tracking in general, just to be clear. The point you raised is very valid - esp with existing methods).
Am I the only person who prefers well-targeted ads? If I see a targeted ad for a power supply IC on a random website, sometimes that's interesting and useful. On the other hand, Twitter is terrible at targeting and the ads are just annoying. Half are ads for random "as-seen-on-TV" quality products and the other half inexplicably think I am an oncologist.
I suppose there is nothing against well-targeted ads as long as the user has consented to having their data being used for targeted-advertising purposes.
The current landscape, however, does not prioritize user-privacy and serves ads to unsuspecting users who do not always know the terms they agreed to.
Moving away from the present opt-out to an opt-in model would give the user a say in how they would like to be tracked across the web. This is also what Apple is trying to do with its "nutritional-info" (not fanboying Apple here) labels. A user is free to consume the content they would like, but they have a right to know how their data is being used.
Would that method benefit from Googles existing people database? I'm thinking it would, and maybe disabling current tracking will kick the ladder out from potential competitors!
> that they are not giving people unique identifiers
How far can this data be reversed? If I have a group of ten people can google tell with high reliability that probably three of them are homosexual, two have Alzheimer's and none are pregnant?
Why is it compared to "random user grouping"? I'm probably misunderstanding something, because it doesn't seem hard to be few hundred percent more accurate than randomness.
https://blog.google/products/ads-commerce/a-more-privacy-fir...
Make no mistake - as far as I can tell, this preserves Google's ability to track that you are, say, a new parent because you've searched for baby clothes. What it won't know is which new parent you are. The system is designed to give probabilistic assurances of k-anonymity. But Google will no doubt tune those "cohort" memberships, and the value of "k," to capture the vast majority of current advertiser needs, while still being able to communicate to antitrust inquiries and the public that they are not giving people unique identifiers. If anything, it hurts their competition more than it would hurt them, because it allows them to thread the needle in a privacy-conscious world.