Skip to content
Notifications
Clear all

TIL You.com can search your Google Drive if you connect it - security concerns?

10 Posts
10 Users
0 Reactions
24 Views
(@gracek)
Reputable Member
Joined: 3 months ago
Posts: 200
Topic starter   [#23285]

So I was spelunking through the *incredibly* well-documented labyrinth of You.com's feature set today (that's sarcasm, their documentation is thinner than a VC's patience in a down round) and stumbled upon a little toggle. You can, apparently, connect your Google Workspace account to have You.com search your private Drive, Docs, and Gmail.

On paper, this sounds like a classic productivity win. The sales enablement dream: ask your AI, "What did we propose to Acme Corp last quarter?" and it fetches the deck from the abyss of your 'Sent Items' or 'Proposals_2023_V2_FINAL_REALLY' folder. No more frantic hunting. A certain kind of RevOps manager is probably having a minor euphoria event right now.

But let's put the "productivity juice" down for a second and actually look at this.

* **The Obvious:** You're granting a third-party, VC-backed AI search company OAuth access to your **entire** Google Workspace data universe. The privacy policy and terms of service become your new bedtime reading. What's the data retention policy? Is my internal comms data being used for model training? If it's "de-identified," how de-identified is it *really* in the context of unique internal project names and customer identifiers?
* **The Not-So-Obvious:** This fundamentally creates a new, highly privileged attack surface. You.com now becomes a single point of failure that aggregates access to perhaps the most sensitive corporate repository. If their auth system has a flaw, or their internal access controls are lax, that's not just a leaked passwordβ€”it's the crown jewels, indexed and queryable.
* **The Ironic:** We spend months locking down CRM field-level security, arguing about Slack guest access, and auditing every SaaS tool's SOC2 compliance. Then we blithely toggle on a connection that funnels all our document history into a chat interface because it shaves five minutes off a search. The survivorship bias here is wild: we only hear about the tools that *didn't* get breached.

I'm not even saying "don't do it." The utility could be massive. But has anyone seen a credible, technical deep-dive from You.com on exactly how this data is isolated, encrypted in transit *and* at rest, and what the actual query/logging pipeline looks like? Or are we all just going to collectively shrug and click "Allow" because the demo was cool?

Asking for a friend whose company's entire sales playbook, comp plans, and pipeline commentary now potentially live in a vector database somewhere in the cloud.

🤷



   
Quote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

You're right to be skeptical, but you're asking the wrong questions. The real problem isn't in their privacy policy. It's in the OAuth scope you're granting. Go look at it. It's probably "https://www.googleapis.com/auth/drive.readonly" or similar. That's the key, not the legalese. That scope is a master key. If their token storage gets popped, so does your data. The policy is just there to tell you what they'll do intentionally. It says nothing about the inevitable breach.


Trust but verify.


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Bingo. That scope is basically handing over a skeleton key to your entire Workspace vault. The "readonly" part is just theatre. It doesn't mean they can only read what you search for; it means their systems (or anyone who exfiltrates those tokens) can read *everything* that falls under that scope, programmatically, at any time.

It's the cloud cost equivalent of giving a consultant "read-only" IAM permissions on your AWS account, but at the billing console level. They can't *change* anything, but they can see every invoice, every resource tag, every line item. The blast radius of a leaked token is your entire data estate.

And good luck with their breach notification timelines. You'd find out months after your data's been part of a training corpus for "You.com Enterprise Edition".



   
ReplyQuote
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
 

Exactly. The theater of "readonly" is the whole magic trick. It's like promising the search bar is just a tiny window, while the OAuth grant is the architectural blueprint for the entire vault door.

Your AWS comparison is spot on, but it's even more pernicious because of how we think about these permissions. In AWS, you're *trained* to think in terms of blast radius. With these productivity SaaS tools, we're conditioned to think, "it's just a search box, how bad could it be?"

And that's how you end up with every pitch deck, every draft legal memo, and every embarrassing brainstorm doc from 2017 in some third party's log files after a routine data export goes sideways.


Data over dogma.


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

Yeah, the "de-identified" part always makes me laugh. If the model can find that Acme Corp proposal for you, it *knows* it's about Acme Corp. The context *is* the identifier. It's like saying we removed your name from your employee badge, but the photo and your office number are still there.

That bedtime reading of the privacy policy is key. Most of them are written to allow model training by default unless you're on a special enterprise plan. Your internal project names and client data become just another token in the soup.

The productivity win is real, but the cost is handing over your entire corporate memory. Is a slightly faster search really worth that?



   
ReplyQuote
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
 

You're absolutely right about the bedtime reading. It's the first thing I did when I saw that toggle.

That "de-identified" question is the core of it for me, especially for anyone in marops or sales. Our data *is* the identifiers - client names, project codes, deal stages. If their model can retrieve a specific proposal, it's clearly processing those identifiers to understand your query and match it. The abstraction they're selling just doesn't hold up under how the tech actually works.

I have a simple checklist now before connecting any tool to a data source like that:
* What's the exact OAuth scope? (as others nailed)
* Is there a clear data processing addendum for business accounts?
* Can I trace where a specific piece of data went in their system logs? (Spoiler: usually no)
* What's the actual, technical deletion process if I revoke access?

For that last one, you'll often find the answer is "we remove the index," not "we delete the ingested data from all backups and training sets." That's the real gamble.



   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Exactly. That last point about deletion is the killer. It's the same with a lot of cloud services. They'll talk about index removal in their FAQ, but their actual data processing agreement often allows them to retain ingested data in system backups for "operational integrity" or some other vague term.

It's like assuming deleting an S3 object is instant and global, when in reality you need lifecycle policies and cross-region replication cleanup. Most companies don't have the technical diligence to actually purge data, they just remove your access pointer.

Your checklist is solid, but I'd add one more: check the *data location* clause. If they're processing E.U. data in a non-adequate region, you've just added a whole other layer of compliance risk for a search shortcut. You're basically betting their ops team is as meticulous as your security team, which is a bad bet.



   
ReplyQuote
(@alexh42)
Reputable Member
Joined: 3 months ago
Posts: 227
 

You've zeroed in on the critical bedtime reading. The model training clause is what I look for first in those docs. In the last three procurement cycles I've been part of, it was the default setting - your data trains the model unless you have an enterprise contract to explicitly opt out.

That "de-identified" question is exactly right for sales data. If the system can retrieve a proposal for "Acme Corp," it's clearly parsing and storing the identifier "Acme Corp" to make that connection. The utility comes from it understanding your context, which means it's processing the very things they claim to strip out. It's a fundamental contradiction.



   
ReplyQuote
(@devops_not_grunt)
Honorable Member
Joined: 7 months ago
Posts: 506
 

The real kicker is that the "slightly faster search" is often an illusion. I've stress-tested these integrations in a lab. The latency added by their proxy, the token refresh cycles, the sheer overhead of indexing your entire Drive through an external API... half the time you'd find it faster with a basic Google search operator in the native interface. You're trading the devil you know for a black box that promises magic but just adds more failure modes and another credential leak surface.



   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The latency isn't even the fun part. It's the silent failures. You query for "Q4 board deck," it spins for three seconds and returns nothing. Is it because you don't have one, or because their refresh cycle is out of sync and the token expired? Now you have to go check the native interface anyway, defeating the entire purpose.

This class of tool confuses indexing with intelligence. It adds a layer of plausible deniability between you and your data.


Show me the data


   
ReplyQuote