AFAIK you can kinda "seek" reads in S3 using a range header, WCGW? =D

candiddevmike • yesterday at 12:44 PM • 3 replies • view on HN

Replies

staticassertion • yesterday at 2:11 PM

You can, and it's actually great if you store little "headers" etc to tell you those offsets. Their design doesn't seem super amenable to it because it appears to be one file, but this is why a system that actually intends to scale would break things up. You then cache these headers and, on cache hit, you know "the thing I want is in that chunk of the file, grab it". Throw in bloom filters and now you have a query engine.

Works great for Parquet.

Sirupsen • yesterday at 3:27 PM

Yep! Other than random reads (~p99=200ms on larger ranges), it's essential to get good download performance of a single file. A single (range) request can "only" drive ~500 MB/s, so you need multiple offsets.

https://github.com/sirupsen/napkin-math

UltraSane • yesterday at 2:17 PM

Amazon S3 Select enables SQL queries directly on CSV, JSON, or Apache Parquet objects, allowing retrieval of filtered data subsets to reduce latency and costs

➕ show 1 reply

alt Hacker News

Replies