You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
use copy_file_range() when possible for assembling uploaded chunks #65102
Use the 👍 reaction to show support for this feature.
Avoid commenting unless you have relevant information to add; unnecessary comments create noise for subscribers.
Subscribe to receive notifications about status changes and new comments.
When doing chunked upload it seems Nextcloud server does:
Upload all chunks to temporary storage
Read all uploaded chunks and write to new (assembled) file
rename and move file to intended destination
delete the uploaded chunks
The problem with this approach is that uploading a file of size "X" will require "2X" local storage before the last step above. While it does not hurt upload bandwidth it causes a large amount of unnecessary IO and uploading a large volume of files (that each are larger than the chunk size) can result in a very long time before the upload is reported as successful at the client side.
I.e. uploading 100GB of files means the user needs to wait until the storage has bothread and written the 100GB again after the time it took to upload the 100GB over the network.
Such long running assemblies (I assume) can also create problems with timeouts or max_execution_times depending on the server configuration.
For some filesystems that is the only way to do it, however a fair few implement copy_file_range() in such a way that allows the assembling of chunks be a mostly/only metadata operation, also bypassing second piping trough php. In this case the upload becomes:
Upload all chunks to temporary storage
call copy_file_range() on the chunks in order together with their intended position (offset) on the .part file
rename and move the file to intended destination
remove the chunks
However in this case, step 2-4 have been metadata only operations finishing orders of magnitude faster, not duplicating any data and with minimal io.
What I propose:
Try to use copy_file_range(), if not supported or fails for whatever reason (mostly configuration errors) fall back to the file stream trough php as it is today.
copy_file_range() gives filesystems an opportunity to implement
"copy acceleration" techniques, such as the use of reflinks (i.e.,
two or more inodes that share pointers to the same copy-on-write
disk blocks) or server-side-copy (in the case of NFS).
Using copy_file_range() would make the assembly practically instant in the following filesystems (to my knowledge):
zfs
xfs
btrfs
modern SMB (samba) - zero network traffic after the upload (like the s3 chunking v2)
modern NFS - zero network traffic after the upload (like the s3 chunking v2)
probably more, idunno
Most (all?) linux filesystems support the call in the sense that the kernel falls back to doing a full copy - but that still saves coping all the data from kernel space -> userspace -> php and back.
Caveat_ish_
The Nextcloud chunks need to be aligned to the underlying storage block size (i.e divides evenly by). As long as the chunks are even MiB:s (or larger), the defaults of all the supported filesystems, or reasonably configurable in the case of zfs it will work. There is no point in doing chunks smaller than 1MiB and (I believe) only zfs allows blocks larger than 1MiB - so barring the default inaccessible >1MiB blocks all allowed values align exactly to whole MiB so if the chunk-size does the same no problems arise. Note MiB =/= MB
Anyway, if it fails fall back to as it is done now, the attempt looses no time.
If I'm understanding this correctly my suggestion boils down to initially attempting to do a plain file stream, only using the AssemblyStream as fallback (probably only relevant with server encryption?)? Basically same as done for the S3 storage chunked upload.
I wrote the issue because I thought I observed the assembly taking very long even when on my underlying storage it's a metadata-only operation (and issuing the copy_file_range manually, by cat parts>out, for 100 1GiB files took 12ms)
Tip
Help move this idea forward
When doing chunked upload it seems Nextcloud server does:
The problem with this approach is that uploading a file of size "X" will require "2X" local storage before the last step above. While it does not hurt upload bandwidth it causes a large amount of unnecessary IO and uploading a large volume of files (that each are larger than the chunk size) can result in a very long time before the upload is reported as successful at the client side.
I.e. uploading 100GB of files means the user needs to wait until the storage has both read and written the 100GB again after the time it took to upload the 100GB over the network.
Such long running assemblies (I assume) can also create problems with timeouts or max_execution_times depending on the server configuration.
For some filesystems that is the only way to do it, however a fair few implement copy_file_range() in such a way that allows the assembling of chunks be a mostly/only metadata operation, also bypassing second piping trough php. In this case the upload becomes:
However in this case, step 2-4 have been metadata only operations finishing orders of magnitude faster, not duplicating any data and with minimal io.
What I propose:
Try to use copy_file_range(), if not supported or fails for whatever reason (mostly configuration errors) fall back to the file stream trough php as it is today.
As put by the manpages for copy_file_range()
Using copy_file_range() would make the assembly practically instant in the following filesystems (to my knowledge):
Most (all?) linux filesystems support the call in the sense that the kernel falls back to doing a full copy - but that still saves coping all the data from kernel space -> userspace -> php and back.
Caveat_ish_
The Nextcloud chunks need to be aligned to the underlying storage block size (i.e divides evenly by). As long as the chunks are even MiB:s (or larger), the defaults of all the supported filesystems, or reasonably configurable in the case of zfs it will work. There is no point in doing chunks smaller than 1MiB and (I believe) only zfs allows blocks larger than 1MiB - so barring the default inaccessible >1MiB blocks all allowed values align exactly to whole MiB so if the chunk-size does the same no problems arise. Note MiB =/= MB
Anyway, if it fails fall back to as it is done now, the attempt looses no time.